Lifebit logo
BlogUncategorizedk-Anonymity, l-Diversity and t-Closeness Explained

k-Anonymity, l-Diversity and t-Closeness Explained

A 3D rendering of a neural network with abstract neuron connections in soft colors.
Photo by Google DeepMind on Pexels

k-anonymity, l-diversity, and t-closeness are three privacy models for de-identifying health data. k-anonymity makes every record indistinguishable from at least k-1 others on its quasi-identifiers (such as age, sex, and postal code); l-diversity adds the requirement that each of those groups contains at least l well-represented values of the sensitive attribute; and t-closeness requires the distribution of the sensitive attribute in each group to stay within a distance t of its distribution in the whole dataset.

Why these models matter now

Health data is being reused for research at a scale that makes informal de-identification hard to defend. Under the General Data Protection Regulation (GDPR), data that can still be linked to a person is personal data, and pseudonymized data remains in scope. Under the US Health Insurance Portability and Accountability Act (HIPAA), the Expert Determination method asks a qualified expert to show that re-identification risk is very small, which in practice means using measurable models like these. The European Health Data Space (EHDS) regulation adds further pressure by requiring secondary-use data to be provided in anonymized form wherever that serves the purpose, and in pseudonymized form only when necessary.

The underlying risk is well documented. In work that led to the k-anonymity model, Latanya Sweeney showed that ZIP code, sex, and date of birth were enough to uniquely identify a large majority of the US population. None of those fields is a name or a medical record number, yet together they act as an identifier. That is the core problem these three models address: removing direct identifiers is not the same as removing identifiability.

The three models explained

Key terms first

Direct identifiers are fields such as name, address, or national ID number, which are removed outright. Quasi-identifiers are fields that are harmless alone but identifying in combination, such as age, sex, ethnicity, postal code, or admission date. Sensitive attributes are the values an attacker wants to learn, such as a diagnosis, a test result, or an HIV status. An equivalence class is a group of records that share the same quasi-identifier values after de-identification.

k-anonymity

Formalized by Sweeney in 2002, k-anonymity requires every equivalence class to contain at least k records. If k is 5, any combination of quasi-identifiers in the released data matches at least five people, so an attacker who knows someone’s age band, sex, and region cannot narrow them down to fewer than five records. It is usually achieved through generalization (replacing an exact age with a ten-year band, or a full postal code with its first characters) and suppression (removing records or values that remain unique).

k-anonymity has a known weakness. If all five records in a class share the same diagnosis, the attacker learns that diagnosis without identifying the exact row. This is called a homogeneity attack. Background knowledge can also rule out some values and narrow what remains.

l-diversity

Proposed by Machanavajjhala and colleagues in 2006, l-diversity requires each equivalence class to contain at least l well-represented values of the sensitive attribute. With l set to 3, every group of matching records must include at least three different diagnoses, so knowing someone’s group no longer reveals their condition outright.

l-diversity has limits too. If a class contains “stage III lung cancer”, “stage IV lung cancer”, and “metastatic lung cancer”, it is technically diverse, but an attacker still learns the person has serious lung cancer. This is a similarity attack. And if a condition is rare overall but common within one group, being in that group still raises the attacker’s confidence, which is known as a skewness attack.

t-closeness

Introduced by Li, Li, and Venkatasubramanian in 2007, t-closeness requires the distribution of the sensitive attribute within each equivalence class to be close to its distribution in the full dataset, with closeness measured by a distance metric such as the Earth Mover’s Distance. If 2% of the whole cohort has a condition, no group should show a markedly higher rate. This limits how much an attacker learns from knowing someone’s group at all.

The trade-off is utility. The stricter the model, the more generalization is needed, and the more analytical detail is lost. t-closeness in particular can flatten exactly the associations researchers are trying to study.

Comparing the models

ModelWhat it guaranteesAttack it addressesRemaining weaknessTypical health data use
k-anonymityEach record matches at least k records on quasi-identifiersIdentity disclosure (linking a record to a person)Homogeneity and background knowledge attacksBaseline for releasing record-level extracts and public-use files
l-diversityEach group has at least l well-represented sensitive valuesAttribute disclosure through uniform groupsSimilarity and skewness attacksDatasets with a single, high-stakes sensitive attribute
t-closenessEach group’s sensitive value distribution is within t of the overall distributionAttribute disclosure through skewed groupsSignificant loss of analytical utilitySmall, highly sensitive releases where utility can be sacrificed
Differential privacyOutput changes little whether or not one person is includedInference from aggregate queriesNoise reduces accuracy; privacy budget must be managedAggregate statistics and query systems

Differential privacy is included for comparison because it protects outputs rather than datasets, and it is increasingly used alongside the classic models. Our guide to differential privacy in health research covers it in depth.

The data harmonization angle

These models are usually discussed as privacy techniques, but in practice they depend on harmonized data. Generalization needs hierarchies: a diagnosis code has to roll up to a broader category, a lab value to a clinically meaningful band, a date to a month or year. When data follows a common model such as the OMOP Common Data Model, with standard vocabularies like SNOMED CT and ICD-10, those hierarchies already exist and are consistent across sources. When every site uses local codes, each generalization step becomes a manual, error-prone mapping.

Harmonization also matters for multi-site work. If one hospital bands age in fives and another in tens, combined releases can create classes that are smaller than either site intended. Agreeing the quasi-identifiers and their generalization rules as part of harmonization prevents that.

A federated approach changes the question further. In a federated Trusted Research Environment (TRE), record-level data stays with each custodian and researchers analyze it in place. Data never leaves the source. Privacy models then apply mainly to what is exported: tables, model outputs, and figures that pass through an airlock. Lifebit’s federated TRE applies automated output checks, such as minimum cell counts, before results are released, with human review for anything flagged. This means custodians rarely need to strip a dataset down to meet t-closeness just to let researchers work with it, because the full-detail data never becomes a released file. For more on the export side, see output checking and statistical disclosure control in TREs.

A practical framework for applying these models

  1. Decide whether you need to release record-level data at all. If researchers can analyze inside a TRE, apply disclosure controls to outputs instead.
  2. Identify quasi-identifiers with a realistic attacker in mind. Consider what public or commercially available data could be linked, such as voter rolls, social media, or news reports of rare events.
  3. Harmonize before generalizing. Map codes to standard vocabularies so generalization hierarchies are consistent.
  4. Choose k, and l or t where needed, based on risk. Wider, less controlled releases need stricter settings than releases to vetted researchers under agreements.
  5. Measure utility after transformation. Check that key distributions and associations survive; if they do not, reconsider the release route.
  6. Document the assessment. Record the model, parameters, assumptions, and residual risk, which is what regulators and HIPAA experts expect to see.

Open-source tools such as ARX and the R package sdcMicro implement these models and report both risk and information loss, which makes the assessment reproducible.

Real-world context

Public bodies already use threshold-based controls that echo k-anonymity. The US Centers for Medicare and Medicaid Services, for example, has a long-standing cell size suppression policy that bars publishing counts below 11 from its data. The European Medicines Agency’s guidance on publishing clinical trial data describes quantitative re-identification risk assessment for anonymized clinical study reports. And TREs across the UK apply statistical disclosure control rules to research outputs as part of the Five Safes framework. The common thread is that small groups are where people become identifiable.

Common pitfalls

Treating k-anonymity as sufficient on its own

k-anonymity protects against identity disclosure only. If the sensitive attribute is uniform within groups, it offers no protection for that attribute.

Ignoring longitudinal data

Health records contain sequences of events. A patient’s pattern of visits and dates can be a quasi-identifier in itself, and single-table models struggle with it.

Forgetting multiple releases

Two separately anonymized releases of the same cohort can be combined to shrink equivalence classes. Track what has been released before.

Over-generalizing until the data is useless

If meeting a model destroys the research value, the release route is wrong. Controlled analysis in a TRE is often the better answer.

What to do next

Review your current de-identification process and ask which of the three models it actually meets, and whether record-level releases are still necessary. For many studies, the safer route is to keep full-detail, harmonized data inside a federated TRE and apply disclosure control to outputs. Our explainer on anonymization vs pseudonymization under GDPR is a useful companion when deciding which legal category your data falls into.

Frequently asked questions

What is k-anonymity in simple terms?

k-anonymity means every record in a dataset looks the same as at least k-1 other records when you consider fields like age, sex, and postal code. An attacker using those fields cannot narrow a person down to fewer than k records.

What is the difference between k-anonymity and l-diversity?

k-anonymity protects against identifying which record belongs to a person. l-diversity adds protection for the sensitive value itself, by requiring each group of matching records to contain several different values, so the group does not reveal a diagnosis on its own.

When should t-closeness be used?

t-closeness suits small, highly sensitive releases where attribute disclosure is the main concern and some loss of analytical detail is acceptable. For most research, analyzing data inside a TRE with output checks preserves more utility.

What value of k is typical for health data?

There is no single standard. Values such as 5 or 11 are common reference points, and the right choice depends on who receives the data, under what agreements, and what other data an attacker could link.

Does k-anonymity make data anonymous under GDPR?

Not automatically. GDPR asks whether a person is identifiable using means reasonably likely to be used. k-anonymity can support that assessment, but residual risks such as attribute disclosure and linkage to other releases must also be addressed.

How do federated TREs change de-identification?

In a federated TRE, record-level data stays with the custodian and researchers analyze it in place. Privacy protection then focuses on exported results, which pass through airlock checks, instead of stripping down the dataset before anyone can use it.