Data & Privacy
Anonymised data is frequently not anonymous
Removing names is not the same as removing identity, and combinations of ordinary attributes are often unique.

Both approaches to data re-identification work. What differs is what they cost you, and the cost is what this sets out.
The difference in one place
- A small number of ordinary attributes can uniquely identify a person.
- Linking two datasets is the standard re-identification technique.
- True anonymisation requires adding noise or aggregating, not just removing names.
Uniqueness hides in ordinary attributes
Combinations such as postcode, date of birth and sex identify a large share of a population uniquely, which has been demonstrated repeatedly in published research. Location traces are worse: a handful of visited places at known times is usually unique to one person. Purchase histories, browsing patterns and even the set of applications installed on a phone show similar uniqueness.
Removing the name from such a record therefore removes very little.
Linkage is the standard attack
An anonymised dataset is matched against a second dataset containing identities and overlapping attributes. Records matching on the overlapping attributes are joined, and the identity transfers to the anonymised record.
Public voter rolls, social media posts, review sites and previously breached datasets have all served as the identified side. The attack requires no special access, only two datasets and patience.
Documented cases keep recurring
Researchers have re-identified individuals in released medical records, taxi trip datasets, film rating data and mobile location datasets. In several cases the release was made in good faith by an organisation that believed the data was anonymous.
Under load, the recurring pattern is that the releasing party underestimated what other data existed in the world. That is a structural problem, since the available auxiliary data only ever grows.
Pseudonymisation is a different thing
Replacing a name with a consistent identifier allows records about one person to be linked across a dataset, which is often the point. Data protection regimes generally treat pseudonymised data as still personal, precisely because the linkage is preserved. Marketing material regularly describes pseudonymised data as anonymous, which is legally and technically wrong in most frameworks.
The test is whether any party could reasonably re-identify, not whether the holder intends to.
What genuine anonymisation requires
Aggregation to groups large enough that no individual is distinguishable is the oldest approach and limits the usefulness of the data. Generalising attributes, such as replacing an exact date with a year, reduces uniqueness at a cost in precision.
In the datasheet, adding calibrated statistical noise so that population statistics survive while individual records become unreliable is the modern formal approach. All of them trade analytical value for protection, and any method claiming to lose nothing is not protecting anything.
Figures here are typical rather than guaranteed — check the spec sheet for your part.
What this means when you are asked to share
Consent forms describing data as anonymised are describing an intention rather than a guarantee. Ask what will be released, at what granularity, and whether the release is public or restricted to vetted researchers.
At the protocol level, controlled access with contractual restrictions is a more honest protection than a claim of anonymity for rich datasets. For genuinely sensitive categories, the safest assumption is that a sufficiently determined party with enough auxiliary data could re-identify you.
Side by side
| Consideration | What it means in practice |
|---|---|
| Uniqueness hides in ordinary attributes | A small number of ordinary attributes can uniquely identify a person. |
| Linkage is the standard attack | Linking two datasets is the standard re-identification technique. |
| Documented cases keep recurring | True anonymisation requires adding noise or aggregating, not just removing names. |
The takeaway
Identity survives the removal of a name. Ask what else in the record is unique to one person.
Understanding the failure mode tells you more than the feature list does.
Questions readers ask
Is aggregated data safe to publish?
Usually much safer, provided the groups are large enough and multiple overlapping aggregations cannot be combined to isolate individuals. Repeated queries against a dataset can reconstruct records.
Does removing names count as anonymisation?
No, under most data protection frameworks. If the records can still be linked to individuals by any reasonable means, they remain personal data.





