Privacy-Enhancing Technologies for Health Data Compared

Privacy-enhancing technologies (PETs) for health data are techniques that let organizations analyze sensitive records while limiting what any party can learn about individual patients. The main options are federated analysis, trusted execution environments, secure multi-party computation, homomorphic encryption, differential privacy, synthetic data, and de-identification, and they differ in whether they protect where data lives, data while it is processed, or the results that leave. No single PET covers all three, so health data programs get the strongest protection by combining an architecture that keeps data at the source with output controls and, where needed, cryptographic or hardware protection.
Why PET selection matters now
Health data programs are being asked to do two things at once: open data for research and AI, and prove that doing so does not expose patients. The European Health Data Space (EHDS) regulation requires secondary use to take place in secure processing environments. The General Data Protection Regulation (GDPR) classes health and genetic data as special category data and expects data protection by design. In the UK, the Information Commissioner’s Office has published guidance on PETs, and the Organisation for Economic Co-operation and Development has examined their role in trusted data sharing. In the United States, the Census Bureau’s adoption of differential privacy for the 2020 Census showed that a formal privacy method can run at national scale, and also that its trade-offs become public debates.
Meanwhile, the 2026 UK Biobank incident gave the field a hard lesson. Approved researchers moved data out through normal workflows of a centralized environment. No encryption was defeated. The gap was architectural: data sat in one place, and outputs were not controlled tightly enough. That episode reframed PET selection for many institutional buyers. The first question is no longer “which cryptography is strongest” but “which combination of controls stops the exposures that actually happen.”
The three layers PETs protect
A useful way to compare PETs is by the layer they protect. Mixing up the layers is the most common reason programs buy a technology that does not address their real risk.
Where data lives
Architectural PETs decide whether data is copied. Federated analysis and federated learning keep records at each custodian and send the computation to them. Nothing needs protecting in a central store because there is no central store.
Data in use
Computational PETs protect data while it is processed. Trusted execution environments (TEEs) use hardware isolation so the host cannot read memory. Secure multi-party computation (SMPC) splits data into shares so several parties compute a joint result without any one seeing the inputs. Homomorphic encryption computes directly on ciphertext.
What leaves
Output PETs control what a result reveals. Differential privacy adds calibrated statistical noise so that no single person’s presence can be inferred. Statistical disclosure control, applied through an airlock, checks tables and models before release. Synthetic data replaces real records with generated ones that preserve statistical patterns. De-identification techniques such as pseudonymization and k-anonymity reduce identifiability in the data itself.
Privacy-enhancing technologies compared
| PET | Layer protected | Main strength | Main limitation | Best health use |
|---|---|---|---|---|
| Federated analysis and learning | Where data lives | No central copy; normal computing speed | Needs harmonized data and output control at each site | Multi-site cohorts, national programs, federated AI |
| Trusted execution environments | Data in use | Protects from host and cloud operator with low overhead | Relies on hardware trust; side-channel research | Hardening compute nodes in cloud deployments |
| Secure multi-party computation | Data in use | No single party sees inputs | Network-heavy; complex to operate | Joint statistics between a few institutions |
| Homomorphic encryption | Data in use | Computes on ciphertext without trusted hardware | High computational cost; fixed computations only | Secure aggregation, private lookups |
| Differential privacy | What leaves | Formal, measurable privacy guarantee | Accuracy loss, especially for small groups | Published statistics, dashboards, repeated queries |
| Statistical disclosure control and airlock | What leaves | Catches risky outputs before release | Needs clear rules and consistent review | Every TRE output |
| Synthetic data | What leaves | Shareable for development and training | Fidelity and residual re-identification risk | Code development, teaching, early exploration |
| Pseudonymization and de-identification | The data itself | Simple and widely understood | Pseudonymized data remains personal data under GDPR | Baseline hygiene for any research dataset |
Reading the table by column tells buyers something important. The technologies with the strongest mathematical guarantees (homomorphic encryption, SMPC) have the narrowest practical scope. The technologies with the broadest scope (federation, airlocks) are architectural and procedural. Real programs need both kinds.
How the layers interact
The layers are not independent. A federated design reduces how much data-in-use protection is needed, because each computation runs inside an environment the custodian already controls. Strong output control reduces the pressure on de-identification, because individual records are far less likely to leak through reviewed aggregates. Conversely, a platform that relies only on data-in-use cryptography still has to answer what happens when a legitimate user asks for a table with a count of two. Evaluating PETs as a stack, rather than one at a time, is what separates a defensible design from a collection of point solutions.
The federation angle: start with architecture, then add layers
Lifebit’s position is that federation is the foundation layer, because it removes the largest single risk: a central copy of sensitive data. In a federated Trusted Research Environment (TRE), approved researchers send analyses to each data custodian, data never leaves the source, and only reviewed aggregate results pass through an automated airlock. That design addresses two of the three layers directly. The third, data in use, can then be hardened with TEEs or cryptographic methods where the threat model calls for it.
This ordering matters for cost. Cryptographic PETs are expensive per computation. Applying them to every analysis in a centralized platform multiplies that cost across all work. Applying them only at specific points in a federated design, for example encrypting intermediate statistics as sites send them to a coordinator, keeps costs proportionate to risk.
It also matters for governance. Regulators and data access committees describe outcomes: no unauthorized access, no unreviewed export, full audit. Federation with an airlock produces evidence for all three. Most individual PETs produce evidence for one.
A buyer’s framework for choosing PETs
- Name the threat. List who you are protecting data from: external attackers, the infrastructure operator, other participating institutions, or approved researchers themselves. Each points to a different layer.
- Decide whether data must move. If the use case can be met by sending analysis to data, federation removes whole categories of risk before any other PET is considered.
- Control every output. Whatever else you choose, put statistical disclosure control, and where appropriate differential privacy, on the path out of the environment.
- Add data-in-use protection selectively. Use TEEs where the infrastructure operator is outside your trust boundary. Use SMPC or homomorphic encryption for specific joint computations between parties that do not trust each other.
- Harmonize before you federate. Federated analysis only works if each site’s data uses a common model. Standards such as the OMOP Common Data Model and HL7 FHIR make one analysis run correctly across many sites.
- Test utility honestly. Measure accuracy loss from noise, synthetic generation, or approximation on realistic tasks before committing.
How national programs combine PETs
Public programs rarely rely on one technology. The UK’s move toward TREs as the default route for health data access, promoted by Health Data Research UK and set out in national data strategy, pairs controlled environments with the Five Safes framework and output checking. The US Census experience shows differential privacy protecting published statistics. The EHDS pairs secure processing environments with permit-based access through health data access bodies.
Lifebit customers illustrate the federated foundation in practice. Genomics England runs research access to genomic and clinical data inside a controlled environment, and Canada’s CanPath connects regional cohorts within a national federated infrastructure. In both, the core protection is architectural and procedural: data held under the custodian’s control, approved researchers, reviewed outputs. Additional PETs sit on top of that base rather than replacing it.
Common pitfalls
Buying the strongest math instead of the right layer
A program that encrypts computation but lets approved users download results has protected the wrong layer. Most real health data exposures happen at the output stage.
Treating synthetic data as anonymous by default
Synthetic data can leak information about outliers and rare conditions if generators overfit. It needs its own disclosure assessment.
Ignoring small cohorts
Differential privacy and disclosure control both struggle with rare diseases, where every count is small. Plan for these cases explicitly rather than applying one threshold everywhere.
Confusing a PET with a legal basis
Using a PET does not by itself make processing lawful. Under GDPR, a research project still needs a legal basis and, for special category data, an applicable condition such as scientific research with appropriate safeguards. PETs are part of those safeguards, and they strengthen a data protection impact assessment, but they do not replace approvals from ethics committees or data access committees.
Forgetting operations
Cryptographic protocols need key management, monitoring, and trained staff. A PET that nobody can operate reliably provides less protection in practice than a simpler control run consistently.
What to do next
Map your current environment against the three layers and mark which ones have real controls today. If data is still being centralized, address that first; Lifebit’s guide to what a Trusted Research Environment is sets out the governance baseline. Then review output protection, starting with differential privacy for published statistics, and consider where synthetic data can support development work without exposing real records. For linkage across institutions, privacy-preserving record linkage is its own specialized PET worth evaluating separately.
Frequently asked questions
What are privacy-enhancing technologies in healthcare?
They are techniques that let organizations use sensitive health data for research and AI while limiting what anyone can learn about individual patients. Examples include federated analysis, trusted execution environments, secure multi-party computation, homomorphic encryption, differential privacy, and synthetic data.
Which privacy-enhancing technology is best for health data?
No single PET is best. Federated analysis protects where data lives, cryptographic and hardware methods protect data in use, and differential privacy and disclosure control protect outputs. Strong programs combine an architectural foundation with output control and targeted data-in-use protection.
Is federated learning a privacy-enhancing technology?
Yes. Federated learning is an architectural PET that trains models across sites without moving the underlying records. Model updates can still leak information, so it is usually paired with secure aggregation, differential privacy, or output review.
Are pseudonymized health records anonymous under GDPR?
No. Pseudonymized data remains personal data under GDPR because it can be re-identified with additional information. Anonymization requires that individuals can no longer be identified by any reasonably likely means.
How do PETs relate to Trusted Research Environments?
A Trusted Research Environment is the governed platform in which analysis happens. PETs are components that strengthen specific parts of it, such as protecting compute nodes, securing aggregation between sites, or adding formal privacy guarantees to outputs.
What is the difference between differential privacy and statistical disclosure control?
Differential privacy adds calibrated noise to results and provides a mathematical privacy guarantee. Statistical disclosure control applies rules and review, such as minimum cell counts, to decide whether an output is safe to release. Many environments use both.
