Lifebit logo
BlogUncategorizedPrivacy-Enhancing Technologies for Health Data Compared

Privacy-Enhancing Technologies for Health Data Compared

A captivating abstract art piece featuring intertwined 3D shapes on a blue background.
Photo by Steve A Johnson on Pexels

Privacy-enhancing technologies (PETs) for health data are techniques that let organizations analyze sensitive records while limiting what any party can learn about individual patients. The main options are federated analysis, trusted execution environments, secure multi-party computation, homomorphic encryption, differential privacy, synthetic data, and de-identification, and they differ in whether they protect where data lives, data while it is processed, or the results that leave. No single PET covers all three, so health data programs get the strongest protection by combining an architecture that keeps data at the source with output controls and, where needed, cryptographic or hardware protection.

Why PET selection matters now

Health data programs are being asked to do two things at once: open data for research and AI, and prove that doing so does not expose patients. The European Health Data Space (EHDS) regulation requires secondary use to take place in secure processing environments. The General Data Protection Regulation (GDPR) classes health and genetic data as special category data and expects data protection by design. In the UK, the Information Commissioner’s Office has published guidance on PETs, and the Organisation for Economic Co-operation and Development has examined their role in trusted data sharing. In the United States, the Census Bureau’s adoption of differential privacy for the 2020 Census showed that a formal privacy method can run at national scale, and also that its trade-offs become public debates.

Meanwhile, the 2026 UK Biobank incident gave the field a hard lesson. Approved researchers moved data out through normal workflows of a centralized environment. No encryption was defeated. The gap was architectural: data sat in one place, and outputs were not controlled tightly enough. That episode reframed PET selection for many institutional buyers. The first question is no longer “which cryptography is strongest” but “which combination of controls stops the exposures that actually happen.”

The three layers PETs protect

A useful way to compare PETs is by the layer they protect. Mixing up the layers is the most common reason programs buy a technology that does not address their real risk.

Where data lives

Architectural PETs decide whether data is copied. Federated analysis and federated learning keep records at each custodian and send the computation to them. Nothing needs protecting in a central store because there is no central store.

Data in use

Computational PETs protect data while it is processed. Trusted execution environments (TEEs) use hardware isolation so the host cannot read memory. Secure multi-party computation (SMPC) splits data into shares so several parties compute a joint result without any one seeing the inputs. Homomorphic encryption computes directly on ciphertext.

What leaves

Output PETs control what a result reveals. Differential privacy adds calibrated statistical noise so that no single person’s presence can be inferred. Statistical disclosure control, applied through an airlock, checks tables and models before release. Synthetic data replaces real records with generated ones that preserve statistical patterns. De-identification techniques such as pseudonymization and k-anonymity reduce identifiability in the data itself.

Privacy-enhancing technologies compared

PETLayer protectedMain strengthMain limitationBest health use
Federated analysis and learningWhere data livesNo central copy; normal computing speedNeeds harmonized data and output control at each siteMulti-site cohorts, national programs, federated AI
Trusted execution environmentsData in useProtects from host and cloud operator with low overheadRelies on hardware trust; side-channel researchHardening compute nodes in cloud deployments
Secure multi-party computationData in useNo single party sees inputsNetwork-heavy; complex to operateJoint statistics between a few institutions
Homomorphic encryptionData in useComputes on ciphertext without trusted hardwareHigh computational cost; fixed computations onlySecure aggregation, private lookups
Differential privacyWhat leavesFormal, measurable privacy guaranteeAccuracy loss, especially for small groupsPublished statistics, dashboards, repeated queries
Statistical disclosure control and airlockWhat leavesCatches risky outputs before releaseNeeds clear rules and consistent reviewEvery TRE output
Synthetic dataWhat leavesShareable for development and trainingFidelity and residual re-identification riskCode development, teaching, early exploration
Pseudonymization and de-identificationThe data itselfSimple and widely understoodPseudonymized data remains personal data under GDPRBaseline hygiene for any research dataset

Reading the table by column tells buyers something important. The technologies with the strongest mathematical guarantees (homomorphic encryption, SMPC) have the narrowest practical scope. The technologies with the broadest scope (federation, airlocks) are architectural and procedural. Real programs need both kinds.

How the layers interact

The layers are not independent. A federated design reduces how much data-in-use protection is needed, because each computation runs inside an environment the custodian already controls. Strong output control reduces the pressure on de-identification, because individual records are far less likely to leak through reviewed aggregates. Conversely, a platform that relies only on data-in-use cryptography still has to answer what happens when a legitimate user asks for a table with a count of two. Evaluating PETs as a stack, rather than one at a time, is what separates a defensible design from a collection of point solutions.

The federation angle: start with architecture, then add layers

Lifebit’s position is that federation is the foundation layer, because it removes the largest single risk: a central copy of sensitive data. In a federated Trusted Research Environment (TRE), approved researchers send analyses to each data custodian, data never leaves the source, and only reviewed aggregate results pass through an automated airlock. That design addresses two of the three layers directly. The third, data in use, can then be hardened with TEEs or cryptographic methods where the threat model calls for it.

This ordering matters for cost. Cryptographic PETs are expensive per computation. Applying them to every analysis in a centralized platform multiplies that cost across all work. Applying them only at specific points in a federated design, for example encrypting intermediate statistics as sites send them to a coordinator, keeps costs proportionate to risk.

It also matters for governance. Regulators and data access committees describe outcomes: no unauthorized access, no unreviewed export, full audit. Federation with an airlock produces evidence for all three. Most individual PETs produce evidence for one.

A buyer’s framework for choosing PETs

  1. Name the threat. List who you are protecting data from: external attackers, the infrastructure operator, other participating institutions, or approved researchers themselves. Each points to a different layer.
  2. Decide whether data must move. If the use case can be met by sending analysis to data, federation removes whole categories of risk before any other PET is considered.
  3. Control every output. Whatever else you choose, put statistical disclosure control, and where appropriate differential privacy, on the path out of the environment.
  4. Add data-in-use protection selectively. Use TEEs where the infrastructure operator is outside your trust boundary. Use SMPC or homomorphic encryption for specific joint computations between parties that do not trust each other.
  5. Harmonize before you federate. Federated analysis only works if each site’s data uses a common model. Standards such as the OMOP Common Data Model and HL7 FHIR make one analysis run correctly across many sites.
  6. Test utility honestly. Measure accuracy loss from noise, synthetic generation, or approximation on realistic tasks before committing.

How national programs combine PETs

Public programs rarely rely on one technology. The UK’s move toward TREs as the default route for health data access, promoted by Health Data Research UK and set out in national data strategy, pairs controlled environments with the Five Safes framework and output checking. The US Census experience shows differential privacy protecting published statistics. The EHDS pairs secure processing environments with permit-based access through health data access bodies.

Lifebit customers illustrate the federated foundation in practice. Genomics England runs research access to genomic and clinical data inside a controlled environment, and Canada’s CanPath connects regional cohorts within a national federated infrastructure. In both, the core protection is architectural and procedural: data held under the custodian’s control, approved researchers, reviewed outputs. Additional PETs sit on top of that base rather than replacing it.

Common pitfalls

Buying the strongest math instead of the right layer

A program that encrypts computation but lets approved users download results has protected the wrong layer. Most real health data exposures happen at the output stage.

Treating synthetic data as anonymous by default

Synthetic data can leak information about outliers and rare conditions if generators overfit. It needs its own disclosure assessment.

Ignoring small cohorts

Differential privacy and disclosure control both struggle with rare diseases, where every count is small. Plan for these cases explicitly rather than applying one threshold everywhere.

Confusing a PET with a legal basis

Using a PET does not by itself make processing lawful. Under GDPR, a research project still needs a legal basis and, for special category data, an applicable condition such as scientific research with appropriate safeguards. PETs are part of those safeguards, and they strengthen a data protection impact assessment, but they do not replace approvals from ethics committees or data access committees.

Forgetting operations

Cryptographic protocols need key management, monitoring, and trained staff. A PET that nobody can operate reliably provides less protection in practice than a simpler control run consistently.

What to do next

Map your current environment against the three layers and mark which ones have real controls today. If data is still being centralized, address that first; Lifebit’s guide to what a Trusted Research Environment is sets out the governance baseline. Then review output protection, starting with differential privacy for published statistics, and consider where synthetic data can support development work without exposing real records. For linkage across institutions, privacy-preserving record linkage is its own specialized PET worth evaluating separately.

Frequently asked questions

What are privacy-enhancing technologies in healthcare?

They are techniques that let organizations use sensitive health data for research and AI while limiting what anyone can learn about individual patients. Examples include federated analysis, trusted execution environments, secure multi-party computation, homomorphic encryption, differential privacy, and synthetic data.

Which privacy-enhancing technology is best for health data?

No single PET is best. Federated analysis protects where data lives, cryptographic and hardware methods protect data in use, and differential privacy and disclosure control protect outputs. Strong programs combine an architectural foundation with output control and targeted data-in-use protection.

Is federated learning a privacy-enhancing technology?

Yes. Federated learning is an architectural PET that trains models across sites without moving the underlying records. Model updates can still leak information, so it is usually paired with secure aggregation, differential privacy, or output review.

Are pseudonymized health records anonymous under GDPR?

No. Pseudonymized data remains personal data under GDPR because it can be re-identified with additional information. Anonymization requires that individuals can no longer be identified by any reasonably likely means.

How do PETs relate to Trusted Research Environments?

A Trusted Research Environment is the governed platform in which analysis happens. PETs are components that strengthen specific parts of it, such as protecting compute nodes, securing aggregation between sites, or adding formal privacy guarantees to outputs.

What is the difference between differential privacy and statistical disclosure control?

Differential privacy adds calibrated noise to results and provides a mathematical privacy guarantee. Statistical disclosure control applies rules and review, such as minimum cell counts, to decide whether an output is safe to release. Many environments use both.