An analyst finishes a piece of work and writes down a name. The evidence looked strong. Four attributes lined up, three sources agreed, nothing contradicted anything else. The finding goes into a report, someone acts on it, and the question nobody asked out loud is how likely it is to be wrong.
Open-source work has a measurement problem at its centre. The discipline is good at gathering and weak at pricing what it gathers. A finding arrives with a sense of confidence attached. That sense comes from how the work felt rather than from anything countable. The gap between the two is where wrong names come from.
This article is about three specific ways that gap opens. Each one is structural, which is why good analysts make all three. Each has a remedy that costs nothing but attention.
They also turn out to be the same error wearing three coats. The last section explains why.
The European regulator’s framing is a useful map. Article 29 Working Party Opinion 05/2014 assesses whether data is genuinely anonymous against three risks: singling out, an individual being isolated within a dataset; linkability, two records about the same person being connected; and inference, deducing the value of an attribute from other attributes. Read from the other side of the table, those are the three things an analyst does for a living. They are also the three places the work overstates itself.
A note on this article. Where this article describes mechanics, it draws on investigation methodology rather than on any specific client engagement. No finding, subject or dataset from any engagement appears here, including in redacted form. The illustrative figures in the table below are generic estimates rather than measurements, and describe no real person. Client data is cryptographically deleted within 48 hours of delivery acceptance and is off-limits to editorial use after that. This is the Data Purge Policy applied without exceptions.
Defined terms
Ten terms do the work in this article. None of them are difficult. Most of the trouble in practice comes from using them loosely.
- Bit. One halving of the candidate set. An attribute worth one bit cuts the people it could be in half.
- Quasi-identifier. An attribute that identifies nobody on its own and identifies somebody in combination. Postcode is one. Date of birth is one.
- k-anonymity. A property of a released dataset, not of a person: every record is indistinguishable from at least k−1 others on the quasi-identifiers.
- Deduction. Reasoning that preserves truth. If the premises hold, the conclusion must hold.
- Induction. Generalising from observed cases to unobserved ones.
- Abduction. Inference to the best explanation. It adds information the premises do not contain. It can be wrong even when every step is reasonable.
- Analysis of competing hypotheses. A method that works by seeking evidence which discriminates between explanations, and by trying to refute rather than confirm.
- Record linkage. Deciding whether two records refer to the same entity.
- Match weight. In record linkage, a score for how much one field agreeing raises the odds that two records are the same person.
- Circular reporting. Several apparently independent reports that trace back to one origin.
Failure one, counting attributes instead of weighing them
The instinct when a profile is thin is to add to it. Another username, another photograph, another employer, another address fragment. The working assumption is that more attributes mean more certainty.
They do not. The reason is arithmetic.
What an attribute is worth is how much it narrows the field. Sex narrows a population by roughly half. Year of birth narrows it by something like a factor of 80. A full date of birth narrows it by a factor of tens of thousands. The useful unit is the bit, where one bit means one halving. Bits add while the candidate set divides.
Eight billion people is about 33 bits. That is the whole budget. Thirty-three halvings takes you from everyone alive to one person, and any combination of attributes that reaches 33 bits has, in principle, named somebody.
Peter Eckersley measured this on browsers in 2010. Across 470,161 browsers he put a lower bound of 18.1 bits of entropy on the fingerprint distribution, which he expressed as odds: pick a browser at random and at best only one in 286,777 others will share its fingerprint. Raising two to the power of the entropy gives the odds, so those are one measurement written twice.
Two caveats belong with that number, and Eckersley states the first himself. His sample was self-selected, drawn from people who visited a privacy-testing site. Everyone in it had gone looking for a privacy tool. The second is that 18.1 bits is a lower bound on that distribution rather than a universal figure for browsers.
| Attribute | Roughly how much it narrows | Bits |
|---|---|---|
| Sex | 1 in 2 | 1 |
| Country (of ~200) | 1 in 200 | ~7.6 |
| Year of birth | 1 in 80 | ~6.3 |
| Full date of birth | 1 in 30,000 | ~14.9 |
| Five-digit postcode (US) | 1 in 33,000 | ~15 |
| Employer, mid-size firm | 1 in 500,000 | ~19 |
Those figures assume each attribute is evenly spread across a population and independent of the others. Real attributes are neither, and the first row shows the problem in miniature.
Counting sex as one bit assumes two categories of roughly equal size, and neither half of that assumption holds. Intersex variation is a biological fact the binary does not describe, and a growing number of states issue passports and civil registrations carrying an X or a third-category marker. Those populations are small, which is exactly why the field behaves so differently for the people in them. For someone recorded male or female, sex removes about half the candidates. For someone recorded X, it removes almost everybody. The same field is worth about one bit to most people and something closer to 20 for a few.
The general form of that is worth carrying into any assessment. Discriminating power belongs to the value in front of you, not to the field it sits in, and a table of averages hides precisely the cases where a single attribute does nearly all the work. A common surname and a rare one are not the same kind of evidence, though both are surnames. Read the table as a rough orientation, then price the actual value.
Sweeney’s 87% and Golle’s replication
The best-known statement of this idea is Latanya Sweeney’s. Working with 1990 US census data, she reported that 87% of the United States population, 216 million of 248 million, had characteristics that likely made them unique on {5-digit ZIP, gender, date of birth}. She also reported 53% unique on {place, gender, date of birth}, and 18% at county level.
That 87% has been quoted for over 25 years. It is the figure behind a great deal of privacy policy.
In 2006 Philippe Golle tried to reproduce it. Using 2000 census data he found 63.3% unique on {5-digit ZIP, gender, full date of birth}, and 14.8% at county level. He then ran the same method against the 1990 data Sweeney had used and got 61%.
Golle did not soften this. His paper says he cannot explain the discrepancy because he lacks detail on the earlier study’s data collection and analysis, and that his own method is set out so it can be replicated and verified.
Both figures were carefully produced, and they disagree by 26 points on the same 1990 data. The 87% is the one that spread.
Notice what this does to the practical question. At 87%, postcode plus sex plus date of birth is effectively an identifier. At 63%, it identifies most people and leaves more than a third of the population sitting in a group with somebody else. An analyst working to the first number stops enriching too early and writes a name down. An analyst working to the second keeps going.
Which of the two is closer to the truth is a question for demographers. What matters for an analyst is that a figure can be quoted for 25 years before anyone re-runs it, and that its familiarity had been doing work its evidence could not support.
The stop rule
The arithmetic gives a rule that is easy to apply. Enrichment has two phases. In the first, each attribute cuts the candidate set. In the second, the set is already one person. Each further attribute is a consistency check rather than a narrowing. The two phases feel identical from the inside. They are worth completely different things.
An attribute added during the first phase raises confidence. An attribute added during the second confirms nothing that was not already established, because it would fit the person you have already isolated whether or not that isolation was correct. Piling on agreeable detail after the set has collapsed is the most common way a wrong identification becomes an unshakeable one.
So: estimate the bits before adding the next attribute. If the candidate set is already at or near one, the useful next move is not another confirming attribute. It is an attempt to break the link.
There is a second correction, the one most often skipped. Bits only add when attributes are independent. City and postcode are largely the same information twice. Employer and professional licence overlap heavily. Two attributes that correlate deliver far less than the sum of their separate values. A profile assembled from six overlapping fragments can look like plenty of evidence while carrying the weight of two.
Where individually harmless details combine into something none of them contains is the mosaic effect, which we have written about separately. This section is the same argument with a number attached. The mosaic is measurable, and measuring it tells you when the picture is complete.
Failure two, calling abduction deduction
The discipline has borrowed a word it is not entitled to. Deduction preserves truth. If the premises are true, the conclusion cannot be false. Induction generalises from observed cases and is honest about being probabilistic. What open-source analysts almost always do is neither. It is abduction, or inference to the best explanation: this handle, this timezone and this commit history are best explained by these two accounts belonging to one person.
Abduction adds information that the premises do not contain. That is the point of it. It is also why abduction can fail with every step reasonable and every source accurate. Sherlock Holmes calls his method deduction while practising abduction throughout, a pleasant literary joke that has done real damage to how this work is described.
The consequence matters more than the vocabulary. Because each step adds information rather than preserving it, a chain of inferences cannot be made certain by lengthening it. Confidence in the whole is bounded by the weakest link in the chain, not raised by the number of links. Our own alias-correlation methodology states this as a rule, and states it well: the confidence of the overall attribution is constrained by the weakest confirmed link, not by how many signals point in the same direction.
Five plausible steps do not make a conclusion five times stronger. They make it as strong as the worst step. They also make it feel much stronger than that.
This is also where the line between what was seen and what was concluded has to be held. Evidence standards exist because an inference quietly promoted to a fact is the most common form of overreach in open-source work, and the promotion usually happens in the writing rather than in the analysis.
What raises confidence when the chain cannot
Three practices do the work that lengthening the chain cannot. The first is competing hypotheses. Richards Heuer’s contribution, in Psychology of Intelligence Analysis, was to invert the natural procedure. Instead of assembling support for the leading explanation, list the plausible explanations together and look for evidence that discriminates between them. Evidence consistent with every hypothesis has no diagnostic value however impressive it looks. Most accumulated OSINT evidence is of exactly that kind. A shared timezone is consistent with the same person, with two colleagues and with coincidence.
This is failure one restated in a different vocabulary. Diagnostic value and discriminating power are the same property. An attribute that does not narrow the candidate set does not distinguish between hypotheses either.
The second is falsification. The question that improves an identification is not “what else supports this” but “what would I expect to see if this were wrong, and have I looked”. If two accounts are one person, their activity should not overlap in ways one person cannot manage. Somebody posting from two continents within the hour is a disproof available to anyone who thinks to check for it. Analysts reliably search for confirmation and rarely for contradiction. The second search is the cheaper one.
The third is language that does not inflate. Sherman Kent made this concrete in the sharpest way available. A 1951 national estimate on Yugoslavia concluded that an attack “should be considered a serious possibility”. Kent, who had helped produce it, meant odds of about 65 to 35. When he asked what the phrase had meant to the policy staff who read it, they had understood something considerably lower. When he then asked his own board colleagues what odds they had each had in mind when they agreed the wording, the answers ran from 20:80 to 80:20.
The people who wrote the sentence together did not agree on what it said. That is the argument for attaching a number, or at least a defined term, to every estimative statement.
Related is the base rate, which is the prior prevalence of what you are looking for in the population you are searching. A match that would occur by chance in one person in a thousand is not a one-in-a-thousand finding when your candidate pool is ten thousand people. It is expected ten times over. Tversky and Kahneman showed in 1974 that people systematically neglect this, and identification work is a close to ideal setting for the error, because the pool is usually large and rarely stated.
Failure three, counting correlated sources as independent
The corroboration rule is the backbone of open-source practice. One source can be wrong. Two independent sources are rarely both wrong in the same direction. Nearly every methodology, ours included, is built on it.
The rule is sound. The load is carried entirely by the word independent, and in the data-broker layer that word is far harder to satisfy than it looks.
To see why, it helps to know how those records are built.
How a broker decides two records are the same person
The operation is called record linkage. It has a formal theory that predates the industry. Ivan Fellegi and Alan Sunter set it out in the Journal of the American Statistical Association in 1969. That model is still the backbone of commercial identity resolution.
The idea is simple. For each field two records might agree on, estimate two probabilities: m, the chance the field agrees given the records really are the same person, and u, the chance it agrees given they are not. A common surname agreeing tells you little because u is high. A rare date of birth agreeing tells you a lot because u is tiny. The weight of that field agreeing is the ratio of the two, and taken as a logarithm the weights of the separate fields simply add up.
Take that logarithm in base two and the match weight is denominated in bits. It is the same unit as failure one. Discriminating power and match weight are one quantity described from two directions, which is why an analyst enriching a profile and a broker merging two records are performing the same operation.
The total weight is then compared against two thresholds, producing three outcomes: link, non-link, and a middle band of possible links for human review. Where those thresholds sit is a business decision, not a fact about the world.
Matching in practice comes in two flavours. Deterministic matching requires exact agreement on a key or a rule set, such as the same hashed email address. Probabilistic matching scores weighted agreement across many fields and links above a threshold. Probabilistic does not mean guessing. It means the decision is explicitly a trade between two error rates. The same machinery, run inside a data clean room, is what lets two companies join their records about you without either showing the other a raw identifier.
And that trade is where the analyst and the broker part company.
Why the analyst’s threshold cannot be the broker’s
Consider what each side loses when the match is wrong.
For a broker selling marketing reach, a false merge costs almost nothing. Two people become one record, an advertisement is served slightly wrong, and nobody finds out. A missed match costs a sale. The economics point in one direction: lower the threshold, accept false merges, maximise coverage. Recall is the product.
For an analyst, the losses invert. A false merge contaminates everything downstream, because every subsequent attribute gets attached to a person who was never in the record. The output is a name given to someone who may act on it, and the cost of naming the wrong person is paid by somebody who was never part of the question. What matters is precision. Precision is bought by raising the threshold and accepting that some genuine matches are missed.
The mathematics is identical in both cases. Only the threshold moves, and it moves in opposite directions.
In a 2025 study published in Proceedings on Privacy Enhancing Technologies, researchers had participants review the records that personal-data removal services had surfaced about them. Of those records, 41.1% described the participant, 30.7% described somebody else, and 28.2% could not be determined either way. That sub-group was 25 participants self-assessing across three services, a small base that indicates a direction without settling a magnitude.
Roughly three in ten records surfaced as being about a named person were about somebody else. That is what a recall-optimised threshold produces. It is also the raw material an analyst corroborates against.
When three sources are one source
The corroboration rule assumes sources that can fail independently of each other. The broker layer does not supply those.
People-search platforms are not independent producers of facts about you. They buy, licence, scrape and re-publish from each other and from a smaller set of upstream suppliers. A record that originates once can surface on four sites within a year, with formatting differences that make it look like four observations.
The intelligence world has a name for this: circular reporting, where several apparently independent reports trace back to a single origin. It is rarely deliberate. Nobody set out to manufacture agreement. Agreement is simply what a resale network produces by default.
Our own Mirror methodology has said this in plain terms for as long as it has been published: people-search platforms share data with each other so aggressively that two of them confirming the same address is one data point, not two.
Repetition in that layer is evidence of distribution, and distribution is not correctness. A record can be wrong, propagate to six platforms and present as six-fold corroboration. Propagation is faster and cheaper for wrong records than for right ones, because nothing in the chain checks.
There is a second-order version that catches careful analysts. Fellegi and Sunter’s model assumes the fields being compared are conditionally independent given match status. Real fields are not. Postcode and telephone area code carry overlapping geography. A system treating their joint agreement as two independent confirmations overstates its own weight by construction.
The same failure appears at three levels: correlated attributes inside one record, correlated records inside one supply chain, and correlated sources inside one corroboration rule. It is a single assumption breaking three times, and every confidence claim in this discipline rests on it.
Provenance before corroboration
Before counting two sources as two, ask where each one got it. The natural question is whether they agree. The useful one is where the record came from, and it costs nothing to ask it first.
In practice that means a small number of habits. Prefer sources from different quadrants, because a registry filing and a people-search listing rarely share an upstream, whereas two people-search listings usually do. Treat identical formatting, identical errors and identical stale fields as a signature of common origin, since a typo that appears on three sites is one typo. Treat agreement within a single category as one source until shown otherwise, and put the burden of proof on the claim of independence rather than on the doubt. Where a record cannot be traced to an origin, it can still be reported, but as a single uncorroborated source.
None of this requires tooling. It requires asking a different question in a different order.
What this changes in practice
One question per failure, asked at three points in the work. Before adding an attribute: how many bits do I already have, and does this one narrow the set or merely fit it? If the set is already one person, stop enriching and start testing.
Before writing a conclusion: what are the competing explanations, what evidence would separate them, and what have I looked for that would prove me wrong? If the answer to the last part is nothing, the finding is untested rather than confirmed.
Before counting corroboration: where did each source get this? Two sources that share an upstream are one source. The burden of proof sits with the claim that they are independent.
Where our own scale sits
We publish a confidence scale, which makes this article partly self-directed.
The reporting scale that reaches a client has three tiers. High means two or more independent sources corroborate the same data point. Medium means one source plus indirect corroboration, or two sources whose independence cannot be established. Unverified means a single uncorroborated source, flagged in the report rather than dropped, because a flagged finding a client can deny is more useful than one quietly discarded. Analysts reason on a finer five-level scale while the work is running. It maps onto those three tiers when the report is written.
Everything in this article bears on the word independent in that first tier. The scale is only as good as the provenance check behind it, which is why the Medium tier explicitly covers two sources whose independence cannot be established. That case is common in the broker layer. A scale without a place to put it will quietly promote it to High.
We have also had to correct our own published wording while writing this. Our Mirror methodology set the High threshold at three or more independent sources while its own worked example used two people-search platforms as two of them, which the same article elsewhere says is one data point. The threshold is now two or more independent sources across the methodology. The example now uses two different quadrants.
The rule was right, the example contradicted it, and the contradiction survived publication because nothing in the pipeline compares a worked example against the rule stated below it.
The shape of the problem
The three failures are one failure.
Failure one counts attributes that overlap as though they were independent. Failure two treats a chain of inferences as though each link were independent of the last. Failure three counts sources that share an origin as though they were independent. Every one is an unexamined independence assumption. Every one inflates confidence in the same direction, because dependence always makes evidence weaker than it looks and never stronger.
Which gives one question that covers all three, worth asking of any finding before it goes into a document somebody will act on: what in here am I counting twice?
Sources
- Sweeney, L. “Simple Demographics Often Identify People Uniquely.” Carnegie Mellon University, Data Privacy Working Paper 3, 2000. dataprivacylab.org. Figures derived from 1990 US Census summary data: 87% unique on {5-digit ZIP, gender, date of birth}, 53% on {place, gender, date of birth}, 18% at county level.
- Golle, P. “Revisiting the Uniqueness of Simple Demographics in the US Population.” Proceedings of the 5th ACM Workshop on Privacy in the Electronic Society (WPES ’06), 2006, 77–80. crypto.stanford.edu. Replication on 2000 Census data: 63.3% unique on {gender, 5-digit ZIP, full date of birth} and 14.8% at county level, and 61% when the same method is applied to the 1990 data. The paper states the discrepancy with the earlier study cannot be explained from the information available.
- Eckersley, P. “How Unique Is Your Web Browser?” Privacy Enhancing Technologies Symposium (PETS), 2010. Full paper (PDF). Lower bound of 18.1 bits of entropy across 470,161 browsers. The sample is self-selected from visitors to a privacy-testing site and is described by the author as biased toward privacy-conscious users.
- de Montjoye, Y.-A., Hidalgo, C. A., Verleysen, M., Blondel, V. D. “Unique in the Crowd: The privacy bounds of human mobility.” Scientific Reports 3:1376, 2013. nature.com. Four approximate location-and-time points uniquely identified 95% of individuals in a dataset covering 1.5 million people.
- Narayanan, A., Shmatikov, V. “Robust De-anonymization of Large Sparse Datasets.” IEEE Symposium on Security and Privacy, 2008. arxiv.org. Establishes that high-dimensional sparse data resists anonymisation, because small combinations of attributes tend to be unique.
- Fellegi, I. P., Sunter, A. B. “A Theory for Record Linkage.” Journal of the American Statistical Association 64(328), 1969, 1183–1210. doi:10.1080/01621459.1969.10501049. The match-weight formulation and the two-threshold decision rule. Both estimation methods in the paper assume conditional independence between field comparisons given match status.
- Heuer, R. J. Psychology of Intelligence Analysis. Center for the Study of Intelligence, Central Intelligence Agency, 1999. cia.gov. Source for the analysis of competing hypotheses and for the diagnosticity of evidence.
- Kent, S. “Words of Estimative Probability.” Studies in Intelligence, Central Intelligence Agency; released under the CIA Historical Review Program, 22 September 1993. cia.gov. Source for the 1951 Yugoslavia estimate, Kent’s own 65:35 reading of “serious possibility”, and the 20:80 to 80:20 spread among the board members who agreed the wording.
- Tversky, A., Kahneman, D. “Judgment under Uncertainty: Heuristics and Biases.” Science 185(4157), 1974, 1124–1131. doi:10.1126/science.185.4157.1124. Source for base-rate neglect.
- Article 29 Data Protection Working Party. “Opinion 05/2014 on Anonymisation Techniques” (WP216), adopted 10 April 2014. ec.europa.eu. Defines singling out, linkability and inference as the three risks against which anonymisation is assessed.
- He, J., Snyder, P., Haddadi, H., Bustamante, F. E., Tyson, G. “Measuring the Accuracy and Effectiveness of PII Removal Services.” Proceedings on Privacy Enhancing Technologies 2025(4), 166–182. doi:10.56553/popets-2025-0125. Record-accuracy figures come from a 25-participant sub-group self-assessing records across three services: 41.1% described the participant, 30.7% described someone else, 28.2% undetermined.
- Josephson, J. R., Josephson, S. G. (eds.) Abductive Inference: Computation, Philosophy, Technology. Cambridge University Press, 1994. Standard treatment of abduction as inference to the best explanation.
Frequently Asked Questions
What is the difference between deduction and abduction in OSINT work?
Deduction preserves truth, so a valid deduction from true premises cannot produce a false conclusion. Abduction is inference to the best explanation, and it adds information the premises do not contain. Almost all identity work in open-source research is abduction, which is why a chain of individually reasonable steps can still reach the wrong person, and why confidence is bounded by the weakest link rather than raised by the number of links.
How many pieces of information does it take to identify one person?
About 33 bits, where one bit is one halving of the candidate set, because roughly eight billion people is about 33 halvings. Far fewer attributes are needed within a smaller population. The important correction is that bits only add when attributes are independent, so overlapping details such as city and postcode deliver much less than the sum of their separate values.
What is a quasi-identifier?
An attribute that identifies nobody on its own and identifies somebody in combination with others. Postcode, date of birth and sex are the standard example. Sweeney reported 87% of the US population unique on that combination using 1990 census data, while Golle’s later replication found 63.3% on 2000 data and 61% on the same 1990 data.
Why is corroboration from several data brokers weak evidence?
Because people-search platforms buy, licence and re-publish records from each other and from shared upstream suppliers, so one record can surface on several sites and present as several independent confirmations. That is circular reporting. Repetition in that layer is evidence of distribution rather than of correctness, and provenance has to be checked before corroboration is counted.
What is record linkage and why does it matter to an investigator?
Record linkage is deciding whether two records refer to the same entity. The Fellegi-Sunter model scores it by weighing how much each agreeing field raises the odds of a true match. It matters because commercial matching thresholds are tuned for coverage rather than accuracy. The records an analyst corroborates against were produced by a process that tolerates false merges.