The Breach Age Index: How to Detect Affiliate Lead Fraud
How to detect affiliate lead fraud at source level. The Breach Age Index exposes recycled identity data that passes every standard validation check.
Short answer: bot-submitted leads pass every validation check you own, because the data they carry is valid, and usually real. It was stolen from somebody.
Nobody is checking for this. Breach exposure is not a standard check anywhere in the lead supply chain, so an entire class of signal sits unexamined on data that changes hands for money every day.
You cannot spot it on a single record. It shows up across a whole source, and it reads the same way every time. A breach rate far above anything real traffic produces raises the alarm. The age of those breaches confirms what it is, because old stolen data is the only raw material a bulk operation can actually get, and a supplier cannot make old data young. Breach depth corroborates it. We call that read the Breach Age Index, and once you have seen it a few times it is unmistakable, the way a constant referrer at scale tells you a channel is switched on.
Key takeaways
- Validation checks the data. It never checks the submission. A valid identity, injected by a script, passes cleanly and becomes a billable lead.
- Nobody is checking for this. Breach exposure is not a standard check anywhere in the lead supply chain. Not at the brand, not at the agency, not at the network.
- Stolen data is the rational choice for a fraudster. Fabricating identities at scale is hard work. But the decisive reason is the trail: buy a list legitimately and bot it, and there is an invoice pointing straight at you. Breached data leaves nothing to follow.
- You cannot spot it on a single record. Ever. Around 50% of clean data carries breach exposure, across affiliate, co-registration, PPC and paid social alike. That is what normal digital life looks like.
- The catch is a sequence, not a number. Breach rate raises the alarm. Breach age confirms it. Breach depth corroborates it. Read by source, that is what bot-submitted stolen identities look like at volume, and it recurs across sub-affiliate fraud again and again. We call the read the Breach Age Index.
- It is old data that gives it away. A recent breach is scarce. Old breach data has been dumped and re-dumped for years, so it is cheap, available in the billions, and it is what anybody feeding a bulk lead operation can actually get hold of. Old stolen data is the raw material this market runs on, and a supplier can lower a breach rate but cannot make old data young.
- Clean sources sit at a 3 to 5 year average first breach. This pattern runs at 9 to 11. Alongside a 70%+ breach rate and records appearing in six or more breaches each.
- Where you hold a date of birth, run the age-at-first-breach test. Subtract the DOB from the date of the first breach and you have the age that person supposedly was when their data was stolen. Real files return teenagers and adults. In Sample A, around 13% of records had been breached before their owner was born. The clean control returned zero.
Definition: the Breach Age Index. A source-level read of breach exposure that exposes recycled stolen data being submitted as leads. A high breach rate raises the alarm. Old breach age confirms it, because a supplier can lower a rate but cannot make old data young. High breach depth corroborates it. It is the recurring fingerprint of a supplier running old stolen identities through high-volume CPL forms.
Valid data is not the same as real data
Most lead fraud detection is still built around one question.
Is this data valid?
That question stopped working years ago.
Valid is easy. A fraudster running a script into your lead form doesn’t have to invent anything. They submit a real name, a real email, a real phone number and a real postcode, because they bought them in a breach dump.
The email delivers. The phone connects. The record dedupes cleanly.
It walks through every validation layer you own and lands in the CRM as an accepted, billable lead.
And nothing anywhere in that chain asks whether the data was stolen from somebody. Breach exposure is not a standard check in lead generation. Not at the brand, not at the agency, not at the network. It is simply not on the list, which means an entire class of signal sits unexamined on data that changes hands for money every day.
We spend most of our time on a different question, and it is the one that matters:
Is this data real, and did the human behind it know it happened?
Why stolen data, and not synthetic
Not every fraudulent lead is a stolen identity. Synthetic identities are real: LexisNexis Risk Solutions has estimated around three million of them in UK circulation. Some of what we see is constructed rather than stolen.
But stealing is the rational choice.
Fabricating identities at scale is hard work. Each one has to be built, made internally consistent, and survive whatever checks it meets. Do it at volume and you are running a manufacturing operation. A breached record arrives complete. Name, address, date of birth, email and phone number, all belonging together, because they belonged to a real person.
But the decisive reason is the trail.
Suppose you did it honestly. You buy a list of consumer data, legitimately, then run a script over it to manufacture leads. You have just created a trail. An invoice. A supplier who knows exactly what data you hold. A transaction sitting in somebody’s records. When those leads turn out to be fraudulent, that trail leads directly back to you.
It would make no sense to do it that way. You would get caught.
Breached data has no trail. Billions of records, circulating in the ether, traded and re-traded and dumped and re-dumped, belonging to nobody and traceable to nobody.
That is not a convenient side benefit. It is the whole reason it is the sensible option.
So the claim is not that every bad lead is stolen. It is that stolen data is what a rational fraudster would use, and it is why breach exposure works as a detection method at all.
Why this lives in sub-networks
Across the affiliate sub-networks we test, most show evidence of fraud somewhere in their traffic, anywhere from around 5% of a source’s volume to effectively all of it. It is unusual to find nothing at all.
That is not because affiliates are unusually dishonest people. It is because of what a sub-network structurally is.
Plausible deniability is baked in. Confront a sub-network and they will tell you it was one of their sources. Go to that source and it was one of theirs. And onwards, until there is nobody left to blame and nothing left to check. It is the defence we hear every single time, and it works because it is almost never disprovable.
The opacity is the product. Once traffic enters a sub-network you have lost the trail of where it came from. That is not an unfortunate side effect of a complicated supply chain. It is the point.
A direct publisher cannot do this. Buy from a publisher, find the leads are fraudulent, and you go back to the publisher. There is nowhere for them to go and nobody to blame. Which is precisely why it doesn’t happen there. It happens where the trail disappears.
And underneath it sits a belief the affiliate world is very willing to hold: that somewhere out there are enormous databases of engaged consumers, sitting ready to be emailed an offer.
Very few people ever ask where all that data actually came from. And of the few who do ask, fewer still refuse to believe the first implausible answer they are given.
What we measured, and how
Sample A, the suspect inflow. A UK campaign buying 800,000 to 1 million leads a month through affiliate and co-registration supply. Roughly 40% of monthly records are rejected as duplicates before they even reach acceptance. We pulled 100,000 records at random from a single month of that brand’s accepted data.
This was not an archived file, not a rejected file and not a test file. Every record had been delivered in real time, passed validation, been accepted as a real person, and been paid for. Three suppliers accounted for over 80% of it.
Sample B, the clean control. 27,000 records from suppliers with direct first-party collection and no intermediary resale, including high intent leads feeding straight into sales contact centres.
Sample C, submission-order analysis. A single supplier delivery of 39,434 leads, analysed in the order the leads were actually submitted, using submission timestamps rather than the row order the supplier chose to send the file in. We have since repeated the analysis across other suppliers and campaigns, and the pattern it exposes recurs.
Every email went through a breach database. The breach date we use is the date the breach occurred, not the date it was disclosed.
One caveat, stated plainly. Breach tools can only match against breaches that have been reported and indexed, so a breach match rate is a floor rather than a ceiling. A record that looks clean may simply sit in a breach that hasn’t surfaced yet.
Every figure below understates the problem.
The benchmark holds across every channel we test
We do not only test co-registration. We run this analysis across affiliate traffic, low-quality co-registration, high-quality co-registration (parenting and family sites, for instance), PPC and paid social, along with direct form fills. The numbers keep landing in the same place.
Clean, recently collected data returns a breach rate of around 50% and a Breach Age Index of 3 to 5 years, whichever channel it came from.
The obvious objection is that we are comparing incentivised co-reg against high-intent form fills, and of course they look different. If the clean benchmark only held on one channel, that objection would land. It doesn’t. A well-run co-registration source and a well-run PPC source produce the same breach profile.
So the benchmark is a property of real consumer data, not of an acquisition channel. Which means when a source diverges from it, the channel is not the explanation.
What breach rate alone can and cannot tell you
Once you start checking breach exposure, the obvious first move is to count breaches. Run the file, see what proportion of the emails come back breached, treat a high number as a bad sign.
That will catch the worst of it. A source at 100% breach exposure is obviously wrong, and you don’t need a second number to know it. Nor do you need a big file: 90 breached out of 100 consecutive leads is already very hard to explain.
But 100% is not where suppliers sit. In Sample A, the largest supplier ran at close to 100%, which is ridiculous. The next was around 90%. Another around 70%. And 70% is close enough to the 50% norm that a supplier will argue with you about it all day, and you will not win that argument with a rate alone.
So the rate catches the egregious and misses the rest. Which is a problem, because most of the market is in the middle.
And on a single record it tells you nothing at all. You cannot call fraud at the individual record level. Plenty of people have used the same email for a decade, got caught in a breach eight years ago, and are completely legitimate. Breach data is a distribution test. It works on the shape of a source. It never works on a record.
Rate, depth and age are three different questions
Breach rate asks how much of the source has been breached at all.
Breach depth asks how many times. The average number of distinct breach events per breached record. This is not the same question, and it answers a different one: not whether the data is stolen, but what kind of stolen stock it is. A file at 100% breached with a depth of 1 came out of a single dump, probably a more recent one. A file at 88% with a depth of 8.9 is aged, aggregated data that has been going round and round the ecosystem for years. The rate condemns both. Depth tells you where the stock came from. And a whole file at exactly one breach each is its own tell, because real people’s breach counts scatter, and a file where they don’t has a single origin.
Breach age asks how long ago it started. This is the one that does the work.
Why it is always old data
Of the three tells, age is the one worth understanding, because it explains why the whole pattern exists.
Age matters not because old data is inherently worse. It matters because old breach data is what is actually available.
A recent breach is scarce. It belongs to whoever took it. It gets sold, and resold, but it takes years to become genuinely widely distributed.
Old breach data is everywhere. It has been dumped, traded, aggregated, compiled and re-dumped for years. It is free or close to free, available in bulk, in the billions of records.
So if you are running an operation that needs millions of identities to feed a lead flow, you are not working from last year’s breach. You cannot get it in the volume you need, and you would be paying for it. You are working from what is freely available at scale, and what is freely available at scale is old.
That is why the pattern looks the way it does. The freely available, bulk, no-questions-asked stolen identity data is old, deeply recirculated, and heavily breached. Feed it into a form at volume and you get exactly the signature the Breach Age Index describes: high rate, high depth, old age.
And age is the one part of it a supplier cannot do anything about. They can filter their stock, substitute records, change their story. They cannot make old stolen data young.
Why a very high breach rate is arithmetically impossible on fresh data
Breach matching keys on the email address. A brand new or recently changed email address cannot appear in a breach that predates it.
Not unlikely. Cannot.
Around a fifth of consumer email addresses are new or changed every year, so any genuinely recent feed is always a mixture of vintages. This year’s addresses sit near zero exposure, last year’s at a few per cent, the year before higher again, because each older cohort has had more time to be caught up in something. The young, low-exposure end of that spread always drags the blended rate down, and it caps how high the number can go.
There is a ceiling at the other end too. Even a very old file never reaches 100%, because normal behaviour leaves a permanent tail of addresses that were never exposed at all. Somebody who has held one address for fifteen years and used it for two trusted accounts, neither of them breached, stays at zero however old the file gets.
So a source running at 90% or above is not a slightly worse version of normal. It is a file containing almost no new email addresses at all, and no real audience produces that.
The full version of this argument is in The KYC Data Problem, where a live inflow ran at 97% breach exposure with half the phone numbers dead.
The Breach Age Index
The Breach Age Index is not a single number. It is the name for a pattern we keep seeing, over and over, in high-volume CPL lead generation.
The Breach Age Index is the breach signature of recycled stolen data being submitted as leads. Three tells, read in order, by source. Breach rate raises the alarm. Breach age confirms what the rate means. Breach depth corroborates it. Read across a whole source, it is the fingerprint of a supplier running old stolen identities through your forms.
The three tells are simple to compute, and because they share a single cause, old stolen stock, they travel together.
Breach rate. The percentage of records in the source with at least one breach match. Clean data sits at around 50%. This pattern runs at 70% to 100%.
Breach depth. The average number of separate breaches each breached record appears in. Real people turn up in two or three. This data turns up in six, eight, nine, because it has been dumped and re-dumped and aggregated for years.
Breach age. The average age of each record’s first breach, the very first time that identity appeared anywhere. Clean data averages three to five years. This pattern runs at nine to eleven. Age gives the index its name because it is the tell a supplier cannot fake. They can filter, substitute and change their story. They cannot make old stolen data young.
The benchmarks
| Tell | What it measures | Clean source | The pattern |
|---|---|---|---|
| Breach rate | % of records breached at all | around 50% | 70% to 100% |
| Breach depth | breaches per breached record | 2 to 3 | 6 to 9+ |
| Breach age | years since the first breach | 3 to 5 years | 9 to 11 years |
No single record is a verdict, ever, and the tells are not equals. The rate is what flags a source, and in the middle of the range it can be argued with, which is exactly why it cannot stand alone. The age is what settles the argument, because it rules out the innocent readings of the rate, and the older-audience defence does not survive it, as covered below. Depth is the same story told a third way. Read in that order, by source, it is the fingerprint we see again and again.
Reading it like a referrer
In legitimate traffic, if you keep seeing the same referrer at scale, you know a real channel is switched on. A constant, repeatable source feeding the pipe.
This is the same thing, in reverse. A breach signature this consistent, at this volume, across delivery after delivery, is not a run of bad luck in the data. It is a constant, scalable input switched on behind the traffic. And the input is old stolen data, because that is the raw material that is cheap, plentiful and passes every check.
Real audiences are noisy. Some records are clean, some were breached last year, some a decade ago, because real people’s breach histories are scattered. A file drawn from one aggregated dump is uniform: the same rate, the same age band, lead after lead. A genuine channel delivers scatter. A dump delivers uniformity.
BAI drift: watch a source run out of good data
Don’t calculate BAI once across a whole file.
Calculate it in submission order, in blocks, and plot the trend.
Use submission timestamps, not the row order the supplier sent you. Row order is whatever they sorted it as. Timestamps are what actually happened.
A legitimate source stays roughly flat.
A source drawing down aged stock gets worse as it delivers, because the good data goes first.
We have now seen this across multiple suppliers and multiple campaigns. It is a pattern, not a curiosity. And it shows on two axes at once: the amount of data that is breached, and the age of those breaches.
| Block, in submission order | Breach rate | Breach depth | Breach Age Index |
|---|---|---|---|
| Records 1 to 10,000 | 69.5% | 3.33 | 8.7 yrs |
| Records 10,001 to 20,000 | 70.0% | 3.50 | 8.9 yrs |
| Records 20,001 to 30,000 | 71.6% | 3.78 | 9.0 yrs |
| Records 30,001 to 39,434 | 85.0% | 7.05 | 10.8 yrs |
| Within that: the final 3,434 | 88.0% | 8.90 | 11.2 yrs |
The last row is not a fifth block. It is a zoom into the tail of the fourth one, records 36,001 to 39,434. Which is the point: the deterioration doesn’t stop at the block boundary. It keeps going right to the end of the file.
Every measure worsens. More breached, more heavily breached, older first breach, all the way through.
The whole-file average was 73.9% breach rate and a BAI of 9.4 years. Bad, but a supplier will argue with you about an average.
The curve is a great deal harder to argue with.
The age-at-first-breach test
There is a second check, and wherever you hold a date of birth it is cheap, reliable and very hard to argue with.
Take each record’s date of birth. Take the date of its earliest known breach. Subtract one from the other.
You now have the age that person supposedly was when their data was first stolen.
Note what this does and does not depend on. It does not need the breach to contain a date of birth, and most don’t. You already hold the date of birth, because your own lead form collected it. The breach database gives you the date. That is all it takes.
What a real file looks like
For a genuine consumer source, that age should almost always be an adult, or at least a plausible teenager. People acquire email addresses in their teens and get caught up in breaches from then on.
They do not have email addresses at four years old.
So a source where a meaningful share of its supposed adults were apparently seven, or nine, or eleven when their data was first breached has fabricated dates of birth. Nothing there needs to be literally impossible for the file to be obviously wrong.
And then there is the impossible end of it
In Sample A, around 13% of records had a first breach that predated the person’s stated date of birth entirely.
Not implausible. Impossible. You cannot have your data stolen before you exist.
The clean control returned zero, and it did so even on records with breaches going back twenty years. So this is not an artefact of old data. A real person’s date of birth and their breach history describe one life and cannot contradict each other, however far back the breaches go.
Where they do contradict, something has been assembled from parts.
Why it happens
Most breaches do not carry a date of birth. They carry emails and passwords, sometimes names and phone numbers.
So a script filling a lead form with a stolen identity has no DOB to work with, and the form demands one. It generates one. And a generated date of birth has no relationship whatsoever to the real person whose email is sitting next to it.
The chronology failures are not an accident. They are a by-product of the fraudster having to invent the one field the breach didn’t give them.
The formatting supports this. Across the failing records the date format was uniform, in a way real consumer input never is. Humans typing a birthday produce a mess. A script produces a pattern.
How to use it
Run it as a distribution, like everything else here. Two numbers per source:
Chronology failure rate. The percentage of records where the first breach predates the date of birth. This should be zero. Any figure above zero warrants investigation.
Age-at-first-breach distribution. How old was each person when their data was first stolen? Plot it. A real source is a curve of teenagers and adults. A fabricated one has a tail of children in it.
This sits alongside breach rate and breach age. It does not replace them, because plenty of lead flows never collect a date of birth at all, and you cannot run it if you don’t hold one.
But where you can run it, it is the hardest single check in this article to explain away.
Breach analysis is one layer of two
Breach analysis tells you an identity is suspect. It works at source level, on distributions, and what it does is condemn a supplier.
It does not tell you that a machine put the record there. That is a separate layer, running on completely different signals: IP fraud scoring, device and browser fingerprint, time to submit, journey completeness, subnet reuse. Those work on the individual submission, in real time, and what they do is block a lead.
The two belong together and they catch different things.
Breach analysis condemns the source. Telemetry blocks the lead.
Everything here assumes both are running. The submission fingerprint deserves its own write-up rather than a table at the end of somebody else’s article, so I’ll do that separately.
The three objections, answered
Every signal here has an innocent explanation available on its own. That is why suspect suppliers survive audits: each one gets argued away in isolation.
Here are the three you will actually hear.
“Lots of people get breached. A high breach rate means nothing.”
Correct, on a single record. Meaningless as a defence at source level.
A file at 90% or above contains almost no new email addresses, and real audiences always contain new email addresses. The vintage floor makes it arithmetically impossible for genuinely recent data.
“Our audience is older, so of course the emails are older.”
Yes, older consumers tend to hold older email addresses. But around a fifth of consumer emails change every year, and that churn applies to everyone. Take a random thousand real people of any age and you still get a spread of vintages. An older demographic shifts the mix. It does not eliminate the young end of it.
A file consisting almost entirely of very old, heavily breached addresses is not an older audience. It is a file with no young cohort at all, and no real audience of any demographic produces that.
And there is a version of this that is worse for the supplier, not better. If a file really is nothing but old, breached addresses, one likely explanation is that it is a breach list being emailed, with the responders then sold on as leads.
That is not an audience. It is stolen data being recycled into traffic.
“We’ll just start filtering out the old records.”
They can try. We don’t think it gets them out.
It destroys their volume. Volume is the entire reason a supplier like this exists. Somebody whose commercial pitch is delivering in a day what a clean source takes six months to produce cannot survive binning most of their inventory.
The records they keep are still breached. Suppressing breach age does nothing to breach rate or breach depth, which is exactly why the numbers travel together. Moving all three means throwing away nearly everything they hold.
And whatever backfills the volume is synthetic, which has its own fingerprint.
We have watched this happen. One supplier, challenged on the evidence, saw its breach rate drop from 92.3% to 46.7% within a couple of hours, at a clean break in the file. The replacement records did not look more legitimate. They looked more constructed: heavier single-domain concentration, cleaner first name and surname local parts, fewer digits, less duplication. The data became tidier than real data tends to be, which is its own tell, because real consumer data is messy.
And the behaviour behind the submissions never changed at all.
The identity stock changed. The machine behind it did not.
Where this data ends up
One reason to care about this at the point of collection is what happens to the data afterwards.
A fraudulent lead does not stop at your CRM. It is a real stolen identity, so it passes downstream identity checks too, and it feeds into the same commercial identity graphs and KYC datasets that the rest of the industry relies on to confirm who people are. Stolen data submitted as a lead today becomes part of the record that verifies a stolen identity tomorrow.
The breach signal is also strongest here, at the start, and it degrades as the record moves. Breach exposure lives in the email address and the phone number. Downstream, a commercial identity match strips a record back to a name, an address, a postcode and a date of birth, and the breach signal is gone entirely.
We have written about the downstream half of this problem separately, in The KYC Data Problem. The short version is that a KYC match can end up verifying a person against records that were stolen from that same person. It starts here, in lead generation, which is the one place the evidence is still intact.
Where to run these checks
Post-hoc analysis proves fraud. Real-time lead validation prevents it.
Running breach analysis over a monthly file tells you what you already paid for.
In one case, a sub-network was paid £250,000 before a pre-submission verification page was put in front of it. The analysis was correct. It was also six figures too late.
To actually stop it, the checks have to sit at the point of submission.
- Check breach exposure in real time. Rate, depth, and above all the age of the first breach.
- Run the age-at-first-breach test wherever you hold a date of birth. Reject the impossible, and look hard at the implausible.
- Aggregate by source, continuously. A single record tells you nothing. Track breach rate and BAI per supplier, per delivery, and watch the trend in submission order.
- Score the submission alongside it. IP fraud, device, timing. Breach analysis condemns the source; telemetry blocks the individual lead.
Do it at the point of submission, not on a monthly report. As covered above, the breach signal is strongest while you still hold the whole record, and it only gets paid for if the lead lands. Reject it at the door and the economics of the fraud fall apart.
What to do, depending on where you sit
If you’re a brand, you are almost certainly the last person in the chain to find out. In every case we have worked on, the end client had no visibility of what was happening upstream.
Start asking your agency for source-level breach rate, breach depth and Breach Age Index, rather than aggregate lead counts. Aggregates hide all of it.
If you’re an agency, your preferred publishers are the ones you check least. Long relationships are where this hides. In one case the sub-network in question had been a trusted partner for seven years.
Run source-level analysis on everyone. Start with the ones you would happily vouch for.
If you’re a network, you’re exposed on both sides. You’re buying traffic you can’t fully see and warranting leads you’re selling.
Pre-submission verification protects your payout rates and your client relationships at once, because you catch the bad leads before they leave your side.
Frequently asked questions
What is the Breach Age Index? It is a source-level read of breach exposure that recurs in high-volume CPL lead generation. A high breach rate raises the alarm, old breach age confirms recycled stolen data, and high breach depth corroborates it. It is the fingerprint of a supplier submitting recycled stolen identities. Age is the standout tell and gives the index its name, because a supplier can lower a breach rate but cannot make old data young. Clean sources average a first breach 3 to 5 years old; the pattern runs at 9 to 11, alongside a 70%+ breach rate and records appearing in six or more breaches each.
How do you detect affiliate lead fraud using breach data? At source level, not record level. Calculate three things per supplier: breach rate, breach depth and average breach age, and plot them in submission order rather than as one whole-file average. Read them in order: a high rate flags the source, old first breaches confirm recycled stolen data, and deep recirculation corroborates it. A single record can never be called.
Does a breached email mean a lead is fraudulent? No. On an individual record a breach match is close to meaningless, because most people with a long-standing email address appear in breaches. Breach data is only diagnostic as a distribution across a whole source.
What is a normal breach rate for lead data? Around 50%, and this holds across affiliate, co-registration, PPC and paid social. Rates of 70% to 100%, especially combined with a BAI above 8 years, are more consistent with aged or recycled identity stock than with fresh consumer intent.
Why does breach age matter more than breach rate? Because old breach data is what is actually available. A recent breach is scarce and belongs to whoever took it. Old breach data has been dumped, traded, aggregated and re-dumped for years, so it is free, available in the billions of records, and obtainable in the volume a bulk lead operation needs. Anybody feeding millions of identities into lead forms is working from old data, because that is what they can get.
What is the difference between breach rate and breach depth? Breach rate is how much of a source has been breached at all. Breach depth is how many separate breach events each breached record has been caught up in. Rate judges whether the data is stolen. Depth characterises the stock: a file at 100% breached with a depth of 1 came out of a single dump, probably a recent one, while a file at 88% with a depth of 8.9 is aged, aggregated data that has been circulating for years. The rate condemns both. Depth tells you where the stock came from.
Why can’t a fresh lead source have a very high breach rate? Because breach matching keys on the email address, and a new email address cannot appear in a breach that predates it. Around a fifth of consumer emails are new or changed each year, so any genuinely recent feed always contains a young, low-exposure cohort that drags the blended rate down. There is a ceiling at the other end too: even very old files never reach 100%, because some people are never exposed.
Couldn’t a high breach rate just mean an older audience? No. Older consumers do hold older email addresses, but roughly a fifth of addresses change every year regardless of demographic, so a sample of real people of any age still produces a spread of vintages. An older audience shifts the mix. It does not remove the young end of it. A file with almost no new addresses in it is not an older audience.
Can bots pass email and phone validation? Yes, routinely, because the identities they submit are valid, and usually real, having been stolen from somebody. Validation confirms the data exists. It tells you nothing about who submitted it.
Are bot-submitted leads stolen data or synthetic data? Both exist, but stolen is the rational choice and it is what we mostly find. Fabricating an identity takes work. More importantly, buying a list legitimately and running a script over it creates a trail: an invoice, a supplier, a transaction. Breached data leaves no trail, because billions of records circulate untraceably.
Why does affiliate lead fraud concentrate in sub-networks rather than direct publishers? Because a sub-network has deniability built into its structure. Challenge one and it blames a source, which blames another source, until there is nobody left to check. The opacity is not an accident of a complicated supply chain, it is the product. A direct publisher cannot do this: if their leads are fraudulent, you simply go back to them.
What is the age-at-first-breach test? Take a record’s date of birth, take the date of its earliest known breach, and subtract one from the other. That gives you the age the person supposedly was when their data was first stolen. A real consumer file returns teenagers and adults, because people acquire email addresses in their teens. A source with a tail of four- and seven-year-olds in it has fabricated dates of birth. It works wherever you hold a date of birth, and it does not require the breach itself to contain one.
What is the chronology failure rate, and what should it be? The percentage of records in a source where the earliest known breach predates the stated date of birth, which is impossible for a real person. It should be zero, and in our clean control it was, even on records with breaches going back twenty years. We measured around 13% in Sample A.
Isn’t the chronology test just detecting old breaches? No. The clean control contained records with breaches going back two decades and still returned zero failures. The test measures whether the fields in a record belong to the same human being. A real person’s date of birth and breach history describe one life and cannot contradict each other, however far back the breaches go.
Can you detect this on a single lead? No, not with breach data. Breach analysis is a source-level distribution test and cannot condemn an individual record. Submission telemetry, meaning IP fraud score, time to submit and device fingerprint, can be scored on a single lead in real time. Telemetry blocks the individual bad submission. Breach analysis condemns the source.
Is a breach match rate a floor or a ceiling? A floor. Breach tools can only check against breaches that have been reported and indexed, so a record that looks clean may simply sit in a breach that hasn’t surfaced. Every breach figure understates the true exposure.
Verify at the point of submission
Provero is a real-time identity and fraud verification API. It exists to inspect a record before it becomes trusted data.
- Breach Detection. Breach presence, depth, and age of first breach.
- IP Fraud Detection. Fraud score, proxy detection, recent abuse, bot status.
- Email Verification and HLR Phone Verification. Deliverability and live line status.
- Name Validation and format checks. Including date-of-birth plausibility against breach history.
One API call, at the moment the lead is created. Before it enters your CRM, and before it becomes a record you have to pay for.
Sign up to Provero and start checking your lead flow, or see it in action against your own data first.
Read next
- The KYC Data Problem: Why a Match Is Not Proof of a Person. What happens to this data downstream, and why a KYC match can verify a person against records stolen from that same person.
- Breach Detection and IP Fraud Detection: the exact signals each check returns.
- Provero for agencies and for lead generation: how source-level verification fits a business buying and reselling traffic.
Methodology note
Findings are drawn from production lead data analysed by Databowl and Provero.
Sample A: 100,000 records selected at random from a single month of accepted, real-time leads for a UK brand buying 800,000 to 1 million leads per month across affiliate and co-registration supply. Sample B: 27,000 records from suppliers with direct first-party collection and no intermediary resale, plus high intent direct form-fill leads feeding sales contact centres. Sample C: a single supplier delivery of 39,434 leads, analysed in submission-timestamp order, with the same analysis subsequently repeated across other suppliers and campaigns. Benchmark figures are drawn from testing across affiliate, co-registration, PPC and paid social sources.
Breach ages are calculated from the date a breach occurred, not the date it was disclosed. Breach checking can only match against breaches that have been reported and indexed, so all breach figures are a floor rather than a ceiling.
The 13% chronology failure rate is a measured figure from Sample A. The chronology and age-at-first-breach checks are repeatable wherever a date of birth is held, but we make no claim that any particular source will fail them at a comparable rate.
All data was anonymised before publication and no client, supplier or campaign is identified. Where we describe what a pattern is consistent with, we are stating our interpretation of the data, not a finding of fact about any named party.
The bottom line
Valid data is not real data. A valid identity, injected by a script, walks through every validation layer you own and becomes a lead you pay for. Stolen data is the rational choice behind it, because breached records arrive complete and leave no trail back to the fraudster.
No single record can be called. The work is to read whole sources, and the read is the same every time: a breach rate real traffic cannot produce, confirmed by breaches old enough that only recycled stock explains them, and corroborated by the same records turning up in breach after breach. It is old stolen data being fed into forms at scale, and it recurs across sub-affiliate fraud so consistently that, once you have seen it, it reads like a switched-on channel. A supplier can lower a breach rate. Nobody can make old data young.
Run it at the point of submission, before the record becomes something you pay for, and the economics that produce the fraud fall apart.