· Simon Delaney · kyc

The KYC Data Problem: Why a Match Is Not Proof of a Person

Identity verification often relies on commercial data whose true origin is unknown. If that data was breached, recycled or brokered, a KYC match can verify a person against records stolen from that same person. The fix is match plus provenance.

Identity verification often relies on commercial data whose true origin is unknown. If that data was breached, recycled or brokered, a KYC match can verify a person against records stolen from that same person. The fix is match plus provenance.

The data used to verify an identity is increasingly the data that was stolen from it. No provider can fix that, because every provider sits downstream of where the problem begins.

The KYC data problem is that identity verification systems often rely on commercial data whose true origin is unknown. If that data has been breached, recycled, brokered or redistributed, a KYC match can verify a person against records stolen from that same person. The fix is not simply a bigger or more reputable database. It is match plus provenance: evidence of where the matched data came from, how fresh it is, whether it is genuinely independent, and whether it deserves to be trusted at all.

A few days ago I searched for a term I’d been writing about: the KYC data provenance problem. The AI summary at the top of the results explained it back to me and surfaced my own work among its sources. So far, flattering. Then I asked it, in effect, the obvious next question: so what’s the solution?

It immediately recommended a long-established, licensed-data vendor.

I want to be precise, because this is not an accusation about that company. It is a serious, regulated provider, and I’m making no claim about the origin of its data. The point is the instinct. Asked how to fix a problem about where identity data comes from, the machine reached straight for a bigger, more reputable, more “licensed” database. When I pushed back, it acknowledged the tension and walked it back, but the climbdown isn’t the interesting part. Models fold under pressure; that proves nothing. The interesting part was the unprompted first move.

Because that reflex is the KYC data problem, performed live. The belief that you answer a question about data origin by naming a downstream provider, that licensed and expensive means clean, is so deeply set in the market that a model trained on the market reproduces it automatically. I’ve spent years on both ends of the consumer-data chain, and I’ve come to think this single reflex is the thing the industry most needs to unlearn.

This piece is the capstone on the subject. It pulls together the evidence I’ve laid out in detail elsewhere: The Data Provenance Problem in Commercial KYC Data, How Affiliate Lead Generation Fraud Feeds KYC Data Fraud, and Who Is Producing All This Identity Data? It states the problem, and the fix, as plainly as I can.

Key takeaways

  • The KYC data problem is not that identity systems fail to match people. It is that they may match people against recycled, breach-derived or poorly provenanced data, and call it proof.
  • A match is not proof of a person. It is proof that two records line up, and those two records are often the same compromised record arriving by different routes.
  • You cannot buy your way out of a provenance problem with a bigger database, because every provider sits downstream of where the data was created.
  • The fix is match plus provenance: a match should carry its source type, freshness, independence, breach exposure and confidence, not just a binary tick.

Definition: the KYC data problem. The KYC data problem is the use of commercial identity data whose true origin cannot be established (data that is frequently breach-derived, recycled and redistributed through opaque supply chains) as evidence that a person is real and trusted. Because the data carries no reliable provenance, a “match” can confirm an identity against records that were stolen from that identity, while looking like independent verification.

Why the reflex is the problem

Start with what a provenance problem actually is. It is a problem about origin: where a record was first created, by whom, whether a human submitted it, whether they consented. By definition, that question can only be answered at the point the record was born.

Every commercial KYC, IDV and AML provider sits downstream of that point. They buy, license, merge and match data that was created somewhere else, often many hops upstream, by parties they have never met. A licensing contract can transfer responsibility for that data. It cannot manufacture visibility into where it came from. The warranty proves everyone agreed it was clean; it does not prove it is.

So when the answer to “the data might be of unknown origin” is “use a more reputable provider,” the answer hasn’t addressed the question. It has moved it one company along and stopped looking. A bigger, cleaner-looking database is still downstream of origin. It inherits the problem with better paperwork.

This is the category error the AI made, and it’s the one the market makes every day. You cannot buy your way out of a provenance problem with a bigger database. The fix is not a better link in the chain. It is a checkpoint before the chain begins.

Regulated use is not the same thing as proven origin.

What this argument is, and what it isn’t

A fair challenge at this point: don’t mature systems already know a database match is weak on its own? Many do. Serious providers layer document checks, biometric liveness, device intelligence and behavioural signals precisely because a record match was never enough, and some do real supplier auditing and direct collection on top. This is not a claim that all identity verification is broken, or that every provider is the same.

It is a narrower and harder claim. It is about the commercial identity-data layer specifically: the broad consumer datasets and identity graphs used to lift match rates, fill gaps and corroborate a spine, sitting alongside stronger primary and biometric checks. That layer is where provenance goes dark. And no provider, however reputable, fully escapes it, because the moment you are matching against commercial data you did not collect, you have inherited its origin, whatever your contract says.

Provenance at scale is not free, either. Checking origin can add friction, and done badly it can punish exactly the people with thin but genuine footprints. That is an argument for doing it well and early, not for not doing it at all. The alternative on offer is to keep rewarding the most heavily circulated identities, which is the opposite of what the system is for.

What the problem actually is

Here, compressed, is the evidence base. The full mechanics and supplier-level analysis are in the deeper pieces; this is the spine of it.

A match is not proof of a person. Commercial B2C KYC matches against a consistent identity spine: first name, last name, address, postcode, date of birth, sometimes with phone and email layered on top. If that spine appears in a commercial database, it looks like a strong match. But two records with the same spine can have completely different origin stories, and the API usually cannot tell you which it is matching. A match confirms the records line up. It does not confirm a human, a consent event, or a real collection.

Corroboration is not redistribution. The market treats the same identity appearing across multiple databases as corroboration, independent confirmation. But once a record enters the consumer-data ecosystem it is sold, routed, enriched, re-permissioned and redistributed. When it later turns up in three places, that is often one origin arriving by three routes, not three witnesses. In 2022 a journalist used subject access requests to map a single breached record fanning out across hundreds of intermediaries, every trail climbing back to one origin. Duplicate distribution is not independent evidence. And a record that has been redistributed widely enough can end up looking more verifiable than a genuine person with a thinner footprint. That is the exact opposite of what an identity system should reward.

The supply chain has the same shape as affiliate lead-gen fraud, and the fraud passes the same way. Client, tier-one provider, an opaque middle layer where the chain disappears, then the consumer. A bot drops breached data into a sign-up form, spoofs the device and browser at the click layer, and the real-but-stolen personal data sails through field validation. It is then opted into a cascade of co-sponsors, each handing it on under its own privacy policy, until hundreds of companies hold a “legitimate” opt-in to a record the consumer never knowingly submitted. That is data laundering: the fraud happens once at the top, and the chain beneath manufactures the legitimacy. The mechanical tells only show at the execution layer (repeated fingerprints, duplicate records, machine-paced timing), which neither the click-layer checks nor compliance is equipped to see.

Methodology note. We analysed 100,000 accepted, paid-for records from a single month of one UK co-registration campaign, alongside a separate live inflow sampled over recent weeks. For each record we checked email and phone validity, breach exposure, breach age, date of birth chronology, impossible calendar dates, timestamp patterns, duplicate identities and supplier level variance. The records were keyed on email address and phone number rather than full identity, which is the reason breach exposure was visible at all. All data was anonymised before publication and no client is identified. One caveat on breach tools is worth stating plainly: they can only check against breaches that have been reported and indexed, so a breach match rate is a floor rather than a ceiling. A record that looks clean may simply sit in a breach that has not yet surfaced.

Start with a real campaign. One of the largest UK co-registration campaigns I’ve analysed is for a household brand buying 800,000 to 1 million leads a month, opted into IDV, KYC and AML checks for downstream companies. I took 100,000 accepted, paid-for, real-time records from a single month. At field level the data looked valid. Underneath, it did not:

  • The largest supplier had close to a 100% breach match rate.
  • The next was around 90%, and another around 70%, with more signs of synthetic construction.
  • Across more than 80% of the data, breach exposure was extreme.
  • Around 13% of people appeared not to have been born when their data was first breached, which is not possible for a real human record.
  • The breaches were old: an average of around nine years, against around three years in higher-quality sources.

And the volume itself cannot be real. Set the breach rates aside for a moment and just look at the size of the order. To deliver one million opted-in co-registration sign-ups a month at industry-typical funnel rates, you would need to send on the order of 833 million emails a month, to a UK adult population of roughly 54 million. That is about 15 emails to every adult in the country, every month, to feed a single buyer’s target.

An audience that genuinely engaged at that scale would be sold at premium rates, not as rock-bottom co-reg traffic. Take your own match volumes and work backwards to how many real, new people could plausibly sit behind them. The gap is the problem.

A more recent live dataset told the same story, harder.

A vendor can wave away a high breach flag as alarmist. Nobody can explain away February the 30th.

There is a separate reason that 97% cannot stand up, and it has nothing to do with the dates. Because the match is keyed on email address rather than full identity, a brand new or changed address cannot appear in any breach that predates it. That single fact is what makes a 97% rate impossible to reconcile with genuinely recent data.

Any real feed is a mixture of email vintages. This year’s addresses sit at close to zero breach exposure, last year’s at a few per cent, the year before higher again, and so on. Each older cohort has had more time to be caught up in something, so it carries a higher rate. A genuinely recent feed is weighted towards the young, low exposure cohorts, which drags the blended rate down. Around a fifth of consumer email addresses are new or changed every year, so that low exposure layer is always present in current data. It cannot be wished away, and it caps how high the blended figure can go.

The ceiling matters too. Even a very old file never reaches 100%, because normal behaviour leaves a permanent tail of addresses that have never been exposed. Someone who has held one address for fifteen years and only ever used it for a couple of trusted accounts, neither of them breached, stays at zero no matter how old the data is. They never convert. So the highest believable rate for any real consumer file sits some way below total, and 97% is pressed hard against the top of it.

There is a further point that makes 97% more damning, not less. Breach checking tools do not hold data from every breach that has ever happened, only from those that have been reported and indexed. A 97% reported rate is therefore almost certainly 100% in reality. The missing 3% is not clean data, it is simply data whose breach has not yet surfaced in a known dataset. A file that is effectively fully breached cannot be recent, because recent data always carries that layer of un-breached new addresses.

For reference, in our own testing across a large number of real data feeds, genuinely current trusted sources typically land around a 40 to 50% breach baseline. That is what normal looks like. A 97% rate is not a slightly worse version of that, it is a different category of data altogether.

Taken together this is two independent proofs reaching the same verdict. The vintage of a recent feed rules out a high blended rate, and the saturation ceiling rules out a near total rate even for old data. A dataset claiming recency whilst showing 97%, in effect 100%, exposure is not fresh. It is either old or assembled from addresses that were already exposed.

Where could data at that scale, accuracy and freshness actually come from? The strongest primary sources (banks, utilities, telecoms, government, credit reference agencies, the electoral roll) are regulated and not generally available as open commercial inflow. One thing, though, produces the exact identity spine at massive scale, with high accuracy, continuously: data breaches.

The 2023 Electoral Commission breach alone exposed data on around 40 million people. The National Public Data breach surfaced up to 2.9 billion records. And on average there are 443 breach notifications every single day in the EU. That doesn’t prove breach data flows into commercial KYC graphs. It proves the spine needed to match an identity is leaking, constantly, at exactly the scale those graphs require to stay refreshed.

The synthetic-identity story is a partial smokescreen. Synthetic “Frankenstein” identities are real; LexisNexis Risk Solutions has estimated around three million in UK circulation. But why build an identity from scratch when you can steal one that already has a name, address, date of birth, email, phone and history attached? Fabricating a synthetic identity is genuinely hard work: it has to be constructed, made internally consistent, then aged and nurtured until it looks lived-in. Stealing a real one takes none of that. Breached records run into the billions and trade cheaply, and every one of them is a real person whose details are genuine by origin. You don’t have to invent anything. You buy it. So the rational fraudster mostly doesn’t bother fabricating, and that is the point: a fake identity is easier to spot as fake, while a real identity, stolen and recycled through enough commercial paths, starts to look like confirmation. If the industry frames the problem as fake people, the answer becomes fake-people detection, which is useful but incomplete. The bigger problem may be real people’s data, used in ways those people never intended.

How a breach actually figures in, as one signal among several

My client’s first question, when they saw the results, was the right one: how do you define a breach, and how do you know? It’s worth answering at the level of principle, because the breach flag on its own is the most misunderstood number in this whole space.

A breach match, conceptually, means a record’s identifiers (typically the email or phone) appear in known breached datasets. On a single record that tells you almost nothing useful. Plenty of entirely legitimate people sit in old breaches because they’ve used the same email for fifteen years. Breach presence is not, and never should be treated as, proof that an individual record is fraudulent.

It only becomes meaningful as one reading inside a wider battery of checks, run at the point a record enters a system:

  • Is the record even valid? A real, deliverable email, a live number, a plausible format.
  • Does it look human, or does its execution-layer behaviour betray a bot?
  • Is it breach-exposed, and how old are those breaches? Recent exposure is ordinary life; very old exposure at scale is not.
  • Is the chronology possible? Does the breach history predate the person’s date of birth, and do the dates themselves even exist?
  • Is the volume plausible against the size of the audience?
  • Are the sources genuinely independent, or is this one origin redistributed?
  • Is it fresh enough to mean anything?

No single one of those is a verdict. Breach is one signal among several. What tells the story is the pattern across them. A 97% breach rate, with nine-year-old breaches, on dead numbers, with impossible dates and artificial timestamps, is not normal digital life. That’s a manufactured dataset wearing the costume of real people. Reading those signals together, at the door, before a record becomes trusted: that is the work. The point isn’t the breach flag. The point is what the breach flag means once it sits next to everything else.

The perfect hiding place

Here is why this persists, and why it isn’t being taken seriously enough: the commercial identity-data layer is the single best place in the economy to hide stolen data.

Think about where this data sits. It is never marketed. No consumer ever sees it move. Most people have no idea it exists, let alone that their breached details might be inside it, quietly being used to verify other people. It is not inspected the way marketing data is inspected. It is a closed, business-to-business match layer that sits in the background and generates value every time it returns a result. If you set out to design a place to launder stolen identity data and never get caught, you would design something that looks exactly like this.

There is a deeper reason it can’t be caught, and it is built into the check itself. The standard commercial KYC check runs on the identity spine: first name, last name, the first line of an address, postcode, date of birth. That’s it. And there is no way to detect breach exposure from those fields, because the breach signal lives in the email address and the phone number, which the spine does not include. So the most common check in the industry is, by construction, blind to the one thing that would give the game away. You could run a clean, compliant, perfectly standard match against a record stolen from a breach a decade ago and have no way of knowing.

The fields chosen for efficiency are the exact fields that hide the problem.

This isn’t theoretical. The recent dataset I described earlier was only catchable because it arrived with everything attached: the emails, the phone numbers, the timestamps, the surrounding metadata. That is what let us piece it together and see the 97% breach rate, the dead numbers, the impossible dates. Had it been spine only (name, address, postcode, date of birth), we would not have caught a single thing. It would have sailed through looking like a million clean records. Which is exactly why this is the best place to hide stolen data. Nothing could be better.

And the paperwork is designed to keep it that way. The chain runs on contracts, not visibility. Each supplier warrants to the next that the data is clean, relying on the warranty of the supplier above. Nobody has to see the original collection event, because the paper says someone already vouched for it. The data doesn’t have to go anywhere or do anything suspicious. It just sits there, matching, earning, wearing a contract that says it’s fine.

That is the quiet part, and it is the whole point. The commercial KYC data layer may be funding the very thing it exists to fight. Every time a polluted record returns a match, somebody upstream gets paid for supplying it, and the system that is supposed to be stopping fraud has just rewarded the data that fraud produced.

The stolen record is monetised twice: once when it is sold into the ecosystem, and again when it passes the check that was meant to catch it.

The solution isn’t a provider, it’s match plus provenance

So we come back to the reflex, and to the only honest answer to “what’s the fix.”

The fix is not switching to a cleaner-sounding database, because the problem lives upstream of every database. The fix is a change of architecture, and it has a name: match plus provenance. Stop returning naked matches. A match should not say “this person matched.” It should say “this person matched against this class of data, from this type of source, with this freshness, this independence, this provenance, and this confidence.” With that context, a match can be weighted honestly: strong identity evidence, weak commercial corroboration, a risk signal only, or something that should not count as proof at all. That is a far more honest output than a binary green tick.

There is an honest objection here, and it is worth meeting head-on. Compliance teams like the binary green tick because it discharges liability. Follow standard procedure, use a recognised provider, and if it goes wrong you can point at the vendor and say you did what the industry does. A nuanced match, one that says “this matched, but the source is weak,” forces a judgment call, and regulated teams are trained to avoid judgment calls, because judgment carries liability.

But turn it around. A binary tick only protects you until a regulator asks you to reconstruct the decision and you cannot. “We matched against a recognised database” is a thin defence when the next question is “and where did that database get the record?” and the honest answer is “we don’t know.” Match plus provenance is not the weaker liability position. It is the stronger one. A match against directly collected, recent, independent data, with its lineage attached, is a defensible audit trail. A naked tick against data of unknown origin is an audit gap waiting to be found. The provenance is not the risk. The absence of it is.

And the timing is not negotiable. Once a record is through the door and laundered across the chain, it is too late: you cannot un-trust data that hundreds of companies now vouch for. The checkpoint belongs at creation, before the record becomes evidence, not after the ecosystem has legitimised it. It belongs there for a practical reason too: creation is the one moment you still hold the whole record, the email, the phone, the metadata around the submission, before a downstream match strips it back to a spine that can hide almost anything.

This is where Provero sits, and it’s worth being exact about the distinction, because it’s the whole argument. This is not a claim that Provero has a magic database. The argument is the opposite. No database can solve a provenance problem after the fact, because by then the problem is already upstream of it. Provero’s claim is architectural: it inspects records before they become trusted evidence. It is a low-latency checkpoint at the point a record enters a system, screening for validity, human-versus-bot behaviour, breach exposure, freshness and provenance signals before the record is trusted. That is a different position, not a different product in the same position. The reason “the solution isn’t any provider” and “this is what Provero does” don’t contradict each other is that the solution was never a provider. It was always a checkpoint. The instinct to reach for a provider at all, the instinct the AI showed when it named a licensed vendor, is the thing to drop.

You cannot buy your way out of a provenance problem with a bigger database. The fix isn’t a better link in the chain. It’s a checkpoint before the chain begins.

This affects more than onboarding

If the underlying identity data is polluted, everything built on top of it inherits the problem: age verification, fraud scoring, identity intelligence, identifier history, device signals, affordability and behavioural trust scores. A risk signal built on poorly provenanced data still has a model, a dashboard and an API response. It hasn’t escaped the provenance problem. It has abstracted it: the same problem wearing a more sophisticated hat. Age verification is the obvious case: in one dataset, around half of duplicate identities reappearing a year later had a different date of birth. That isn’t a formatting issue. It’s a sign the spine isn’t stable enough to carry evidential weight.

And as AI gets embedded in fraud scoring and identity decisioning, the governance bar only rises. AI doesn’t create identity data; it uses it. If the data layer is breach-derived and recycled, AI isn’t mimicking real people, it’s mimicking a polluted system, and it can’t tell the difference, because polluted records are already part of what it treats as real. The EU AI Act’s direction of travel (documentation, logging, oversight, data governance for high-risk systems) pushes towards stacks that have to explain the data beneath the score, not just the score. The honest question for any KYC stack is: why do you trust the data under the output?

Questions to ask your KYC, IDV or AML provider

  • What proportion of your match data comes from directly collected sources, versus co-reg, affiliate, brokered or partner-supplied data?
  • Do you classify source types in the match response?
  • Do you test suppliers for breach exposure, breach age and impossible chronology?
  • Can you distinguish corroboration from redistribution: genuinely independent sources from the same record arriving twice?
  • Do you return provenance, freshness and source confidence alongside the match?
  • Can a customer downweight or exclude weaker source classes?
  • What happens when a source produces high match rates but poor provenance signals?
  • Are you measuring only synthetic-identity risk, or also the risk of real identities being breached, recycled and falsely corroborated?

The answers don’t need to be perfect. But a provider who cannot answer them at all is asking you to run KYC blind.

The real question

Commercial KYC has spent years optimising for coverage and match rate. The next phase has to optimise for provenance, independence and trust. The risk was never that these systems fail to find a match. The risk is that they find a match against data that should never have counted as proof, and that, by rewarding the identities that have circulated most widely, the market quietly rewards the stolen ones.

The question is no longer does this data match? It is should this match be trusted? And underneath even that, there’s the question the AI’s reflex should make all of us ask:

If your answer to a data-origin problem is to go shopping for a provider, are you funding the very thing you’re supposed to be fighting?

FAQ

What is the KYC data problem? The KYC data problem is the use of commercial identity data whose true origin cannot be established (data that is often breach-derived, recycled and redistributed through opaque chains) as evidence that a person is real and trusted. Because the data carries no reliable provenance, a match can confirm an identity against records that were stolen from that identity, while appearing to be independent verification.

Is the solution to switch to a bigger or more reputable KYC provider? No. A provenance problem is a problem about where data originated, and every commercial provider sits downstream of origin. A cleaner-looking, more heavily licensed database inherits the problem with better paperwork rather than solving it. You cannot buy your way out of a provenance problem with a bigger database.

Can licensed data fix it? A licensing contract transfers responsibility and proves everyone in the chain agreed the data was clean. It does not prove where the underlying record was first created, by whom, or whether a human consented. Regulated use is not the same as proven origin.

Doesn’t layering biometrics, liveness and document checks already solve this? Those checks help, and serious providers use them precisely because a database match was never enough on its own. But this argument is about the commercial identity-data layer that sits alongside them to lift match rates and corroborate a spine. That layer is where provenance goes dark, and no amount of biometric checking downstream tells you where a commercial record was originally created.

Isn’t a binary match safer for a compliance team? It feels safer, because it discharges liability: follow standard procedure and you can point at the vendor if it goes wrong. But a binary tick only protects you until a regulator asks you to reconstruct the decision. A match with its source, freshness and independence attached is a more defensible audit trail than a naked tick against data of unknown origin. The provenance is not the risk; the absence of it is.

Why is the commercial KYC data layer so hard to police? Because it is invisible and contractual. It is never marketed, no consumer sees it, and it is rarely inspected the way marketing data is. The chain runs on warranties rather than visibility, each supplier relying on the one above. Stolen data doesn’t have to move or behave suspiciously to stay hidden; it just sits in the match layer, earning every time it returns a result. That makes it close to a perfect place to hide breached data.

Why can’t a standard KYC check catch breach data? Because most commercial checks run on the identity spine: first name, last name, address line one, postcode and date of birth. Breach exposure can only be seen in the email address and the phone number, which the spine does not include. A standard, compliant match can therefore confirm a record that was stolen in a breach years ago without ever surfacing a breach signal. To detect it you need the fuller record, which is one more reason to run the check at the point of entry, where that data still exists.

How do you detect breach data in identity records? A breach match means a record’s identifiers appear in known breached datasets. On a single record that proves nothing, as many legitimate people sit in old breaches. It only becomes meaningful as one signal inside a wider set of checks run at the point of entry: validity, human-versus-bot behaviour, breach exposure, breach age, chronology, volume plausibility and source independence. The signal is the pattern across all of them, not the breach flag alone.

Is breach presence proof that a record is fraudulent? No. Breach presence is a pattern-level signal across suppliers and cohorts, not individual proof. It becomes telling when breach rates and breach age are extreme, when chronology is impossible, or when it sits alongside other signs of synthetic or recycled construction.

What is “match plus provenance”? It means a KYC or IDV response should not simply say a person matched. It should return the class of data matched against, the source type, freshness, independence, breach exposure, collection context and a confidence weighting, so a match can be read as strong evidence, weak corroboration, a risk signal, or not proof at all.

What’s the difference between corroboration and redistribution? Corroboration means an identity is independently confirmed by two or more separate original sources. Redistribution means the same original record has been sold and routed through multiple intermediaries and now appears in several places. They look identical inside a commercial database, but only the first is real evidence. True corroboration requires source independence.

Where should identity data be checked? At the point a record enters a system, before it becomes trusted. Once a record is through the door and laundered across the chain, it’s too late: you cannot un-trust data that hundreds of companies now vouch for.

Does the KYC data problem only affect onboarding? No. Anything built on commercial identity data inherits it: age verification, fraud scoring, identity intelligence, identifier and device history, risk and affordability signals. A sophisticated-looking score built on poorly provenanced data hasn’t escaped the problem; it has abstracted it.

Are stolen identities or synthetic identities the bigger KYC problem? Both are real, but stolen identities may be the larger and quieter problem. A synthetic identity has to be fabricated, made consistent and nurtured over time. A stolen identity needs none of that, because it is a real person’s genuine data, and breached records are available cheaply in their billions. The market tends to frame the enemy as fabricated people, because that fits a clean fraud narrative, but the easier and more convincing route is usually to use a real identity that was never yours to use.

Why might a fraudster look more verifiable than a real person? Because matching rewards records that appear across many sources, and stolen identity data tends to appear across many sources by the time it has been sold and redistributed. A widely circulated stolen identity can look more verifiable than a genuine person with a thinner commercial footprint, the opposite of what a KYC system should reward. It is also why stolen identities, rather than synthetic ones, are the likelier tool: a synthetic identity has to be built and nurtured, while a stolen one is a real person’s genuine details, available cheaply and in vast quantity. Why fabricate an identity when billions of real ones are already for sale?

Sources and further reading

  • UK Finance, Annual Fraud Report 2025 (covering 2024 data): 24,407 confirmed third-party card application fraud cases, £19.3m.
  • LexisNexis Risk Solutions, research on synthetic “Frankenstein” identities in UK circulation.
  • EU AI Act, Articles 10, 12, 13, 14 and 15: data governance, record-keeping, transparency, human oversight, accuracy and robustness.
  • GDPR and ICO guidance on lawful basis, legitimate interests and fraud prevention.
  • Publicly reported UK and international data breaches, 2017 to 2025.
Share:

Related Articles

View All Articles »