Anonymisation vs Pseudonymisation Under GDPR


Under the General Data Protection Regulation (GDPR), anonymisation and pseudonymisation are legally distinct concepts with opposite consequences. Anonymised data no longer relates to an identifiable person, so the GDPR ceases to apply to it entirely; pseudonymised data has had direct identifiers replaced with codes but remains personal data, because re-identification is still possible using additional information held separately. For health data — and genomic data in particular — the practical reality is that almost everything researchers call “anonymised” is, in GDPR terms, pseudonymised, and must be governed accordingly.
Why the distinction matters now
The gap between the two terms has become one of the most consequential compliance questions in health research. The European Data Protection Board (EDPB) adopted its Guidelines 01/2025 on pseudonymisation in January 2025, confirming that pseudonymised data remains personal data in the hands of the controller and clarifying when it may cease to be personal data for a specific recipient. The Court of Justice of the European Union (CJEU) has repeatedly examined the question — from Breyer (C-582/14), which established that identifiability depends on the means “reasonably likely” to be used, through the EDPS v SRB litigation on whether pseudonymised data transferred to a third party remains personal from that recipient’s perspective.
Layered on top is the European Health Data Space (EHDS) regulation, which entered into force in March 2025 and builds its secondary-use regime around pseudonymised — not anonymised — health data processed inside secure processing environments. Regulators have accepted what statisticians have argued for two decades: rich health data cannot be meaningfully anonymised without destroying its research value, so the governance question shifts from “how do we anonymise?” to “how do we control access to pseudonymised data?”. That shift is the reason Trusted Research Environments (TREs) exist.
The GDPR definitions, precisely
Anonymisation: Recital 26
The GDPR does not define anonymisation in an operative article. Recital 26 states that the principles of data protection do not apply to “anonymous information”, meaning information which does not relate to an identified or identifiable natural person, or personal data “rendered anonymous in such a manner that the data subject is not or no longer identifiable”. The test is whether identification is possible by “means reasonably likely to be used” by the controller or by another person — accounting for cost, time, available technology, and future technological development. This is a high bar: anonymisation must be robust against realistic attackers, not merely against casual inspection, and it must remain robust as auxiliary datasets and re-identification techniques improve.
Pseudonymisation: Article 4(5)
Pseudonymisation, by contrast, is explicitly defined in Article 4(5): processing personal data so it “can no longer be attributed to a specific data subject without the use of additional information”, provided that additional information is kept separately and protected by technical and organisational measures. Replacing an NHS number with a study ID, hashing an email address, or tokenising a medical record number are all pseudonymisation. Crucially, Recital 26 states in terms that pseudonymised data “should be considered to be information on an identifiable natural person”. Pseudonymisation is a safeguard the GDPR actively encourages — it appears in Article 25 (data protection by design) and Article 89 (research safeguards) — but it is not an exit from the regulation.
Health data as special category data
Both concepts interact with Article 9, which classifies data concerning health and genetic data as special category data, prohibited from processing unless a specific condition applies — for research, typically Article 9(2)(j) with Article 89(1) safeguards. Pseudonymised health data is still special category data. Only genuine anonymisation removes it from Article 9’s scope, and for the data types below, genuine anonymisation is rarely achievable.
Why health data resists true anonymisation
Health and genomic data are structurally hostile to anonymisation. Latanya Sweeney’s foundational work showed that 87% of the United States population is uniquely identifiable from just ZIP code, birth date, and sex — three fields present in almost every clinical dataset. Gymrek and colleagues demonstrated in Science (2013) that surnames can be recovered from Y-chromosome short tandem repeats in “de-identified” genomes by cross-referencing public genealogy databases. A genome is not merely linked to an identifier; it is an identifier — stable for life, shared probabilistically with relatives, and impossible to revoke.
Longitudinal records compound the problem. A sequence of hospital admissions with dates and diagnoses forms a fingerprint even after every direct identifier is stripped. Rare-disease cohorts are more exposed still: when a condition affects a handful of people in a country, the diagnosis alone narrows identity to near-uniqueness. This is why the UK Information Commissioner’s Office (ICO), in its anonymisation guidance, applies a “motivated intruder” test rather than accepting the removal of names as sufficient — and why data that passes the United States HIPAA Safe Harbor de-identification standard (removal of 18 listed identifiers) frequently still fails the GDPR’s Recital 26 test. The two regimes are not equivalent, and treating them as interchangeable is a common and costly compliance error.
The data harmonisation dimension
There is a second, less discussed reason the distinction matters: data harmonisation. Health research increasingly runs across multiple cohorts, and harmonising them to a common data model — such as the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM), maintained by the Observational Health Data Sciences and Informatics (OHDSI) community — makes records more comparable and therefore, paradoxically, more linkable. A harmonised dataset that maps local codes to standard vocabularies and aligns dates, demographics, and measurements is easier to join against auxiliary data than the messy source it came from. Data harmonisation raises research value and re-identification surface at the same time.
The correct conclusion is not to avoid harmonisation — unharmonised data produces unreliable science — but to accept that harmonised health data is pseudonymised data, and to govern it inside an environment designed for pseudonymised data. Data harmonisation and access control are complements: the first makes data usable, the second makes using it lawful. Lifebit’s approach treats harmonisation as an in-place operation, running AI-assisted OMOP and Fast Healthcare Interoperability Resources (FHIR) mapping where the data lives, so the harmonised output inherits the custodian’s controls rather than escaping them.
Anonymisation vs pseudonymisation: side-by-side
| Dimension | Anonymisation | Pseudonymisation |
|---|---|---|
| GDPR status | Not personal data; GDPR does not apply (Recital 26) | Personal data; GDPR fully applies (Article 4(5), Recital 26) |
| Legal test | No identification by means reasonably likely to be used, by anyone | Attribution possible only with separately held additional information |
| Reversibility | Must be irreversible in practice | Reversible by design, under controlled conditions |
| Research utility | Low for rich data — aggressive generalisation and suppression required | High — record-level detail and longitudinal structure preserved |
| Feasibility for genomic data | Effectively unachievable at record level | Standard practice |
| Role in GDPR | Exit from scope, rarely attainable for health data | Named safeguard under Articles 25 and 89 |
| Typical governance | Open or lightly controlled release | Trusted Research Environment with access controls and output checking |
Governing pseudonymised data: the TRE model in practice
Once an organisation accepts that its research data is pseudonymised rather than anonymised, the governance model follows. A federated Trusted Research Environment keeps pseudonymised data at its source — the biobank, hospital, or national programme that collected it — and brings approved researchers’ analyses to it, so the data never leaves the source. Access is governed under the Five Safes framework (safe people, safe projects, safe settings, safe data, safe outputs), developed at the UK Office for National Statistics, and every result leaving the environment passes disclosure control before release. This is precisely the architecture the EHDS mandates for secondary use: pseudonymised data, secure processing environment, controlled outputs.
Genomics England operates on this model. Its research environment holds pseudonymised whole genomes linked to longitudinal NHS clinical records — data that could never satisfy Recital 26 — and makes them available to thousands of approved researchers who analyse in place and export only vetted, aggregate results. The lawful basis rests not on a claim of anonymity but on demonstrable safeguards around pseudonymised data: exactly the posture the EDPB’s 2025 guidelines reward. The same logic underpins TRE compliance across HIPAA, GDPR, and EHDS: one architecture, multiple regimes satisfied.
Common pitfalls
Three errors recur in health-data programmes. First, scope denial: treating key-coded data as anonymous because the research team cannot see the key. Under the Recital 26 “means reasonably likely” test — and in the controller’s hands under the EDPB guidelines — that data is pseudonymous, and every GDPR obligation still attaches, from lawful basis to breach notification. Second, release-and-regret: publishing “anonymised” record-level extracts that are later re-identified against auxiliary data. Because the Recital 26 test accounts for future technology, an extract that is defensible today may not be defensible in five years — and once released, it cannot be recalled. Third, regime conflation: assuming HIPAA-de-identified data is GDPR-anonymous, which fails the moment a dataset crosses into EU or UK jurisdiction. Each pitfall shares a root cause — putting the compliance weight on a property of the dataset rather than on the environment that controls it.
What to do next
Audit every dataset your organisation currently labels “anonymised” against the Recital 26 test, honestly applied: could a motivated party, using reasonably available means and auxiliary data, re-identify anyone? For most record-level health data the answer is yes — reclassify it as pseudonymised and bring it under full GDPR governance. Keep pseudonymisation keys separated with independent access controls, document the technical and organisational measures around them, and move record-level analysis into a Trusted Research Environment governed by the Five Safes framework, reserving open release for genuinely aggregate, disclosure-checked outputs. Treat data harmonisation as an in-place operation inside that environment. The organisations that get this right stop arguing about whether their data is anonymous — and start proving that it does not need to be.
Frequently asked questions
What is the difference between anonymisation and pseudonymisation under GDPR?
Anonymised data cannot be linked back to an individual by any means reasonably likely to be used, so the GDPR no longer applies to it. Pseudonymised data has had identifiers replaced with codes but can be re-linked using separately held information, so it remains personal data and the GDPR applies in full.
Is pseudonymised health data still personal data?
Yes. GDPR Recital 26 states that pseudonymised data should be considered information on an identifiable natural person, and the EDPB’s Guidelines 01/2025 confirm it remains personal data in the controller’s hands. Because it concerns health, it also remains special category data under Article 9.
Can genomic data ever be truly anonymised?
At record level, effectively no. A genome is itself a stable, lifelong identifier, and published research has shown individuals can be re-identified from de-identified genomes using public genealogy databases. Genomic research therefore relies on pseudonymisation combined with controlled environments rather than claims of anonymity.
Does HIPAA de-identification count as GDPR anonymisation?
Not reliably. HIPAA Safe Harbor requires removing 18 listed identifiers, but data that passes it can still be re-identifiable by the GDPR’s “means reasonably likely” standard. Organisations operating across both regimes should assess data against each test separately.
Why does the EHDS use pseudonymised rather than anonymised data?
The European Health Data Space regulation requires secondary use of health data in secure processing environments, with anonymised data preferred but pseudonymised data permitted where anonymisation would defeat the research purpose. Because rich clinical and genomic data loses its value when anonymised, pseudonymisation within a controlled environment is the operative model.
Is pseudonymisation enough on its own to comply with GDPR?
No. Pseudonymisation is a named safeguard under Articles 25 and 89, not a complete compliance strategy. Controllers still need a lawful basis, an Article 9 condition for health data, security measures, and — for research at scale — an access-controlled environment with output checking.
How do Trusted Research Environments handle pseudonymised data?
A Trusted Research Environment (TRE) keeps pseudonymised data inside a controlled setting, admits vetted researchers to analyse it in place, and applies disclosure control to everything that leaves. In a federated TRE the data additionally never leaves its custodian, which strengthens both GDPR accountability and data-sovereignty compliance.
