Build vs Buy: Standing Up a National TRE


Standing up a national Trusted Research Environment (TRE) is not a binary build-versus-buy decision: the programmes that work treat the governance layer as sovereign and non-delegable, and the platform layer as infrastructure to be procured — because building a secure, federated analytics platform from scratch consumes three to five years and scarce specialist engineering before the first researcher logs in, while a procured federated platform deployed inside government-controlled infrastructure delivers the same sovereignty in months. The decisive evaluation criterion is not build or buy but custody: whoever supplies the software, the data must remain within national jurisdiction and control, and the data never leaves the source.
Why this matters now
National TREs have moved from aspiration to statutory obligation. The European Health Data Space (EHDS, Regulation (EU) 2025/327) requires that secondary use of electronic health data occur only within secure processing environments — Article 50’s term for a TRE — which every member state must now be able to provide or certify. The UK reached the same conclusion by review rather than regulation: the Goldacre Review (2022) recommended that national research access shift from data dissemination to accredited secure environments, and NHS England’s subsequent Secure Data Environment policy made that the default for NHS data. Momentum, however, has a counterweight: the May 2026 UK Biobank incident, in which approved researchers exported data through a centralised platform’s normal workflow, sharpened ministerial scrutiny of which architecture a national TRE runs on — not just whether one exists. Procurement teams are therefore answering two questions at once: who supplies the platform, and does its architecture make bulk egress impossible rather than merely forbidden.
The federated TRE angle: sovereignty is an architecture, not a procurement clause
The instinct behind “build” is usually sovereignty — a belief that only software written in-house keeps national data under national control. That conflates two separable things. Sovereignty is determined by where data resides, who operates the infrastructure, and what can leave; it is not determined by who wrote the code. A federated Trusted Research Environment deployed inside government cloud or on-premises infrastructure keeps every record within national jurisdiction, under national operational control, with analysis travelling to the data and only disclosure-checked aggregates leaving through an automated airlock. Conversely, a nationally built platform that centralises records into one repository has weakened sovereignty in the way that matters — it has created a single point of custody, breach, and political risk. Federation also changes the national-scale economics: because data stays with each custodian (regional health authorities, hospitals, biobanks, registries), the programme does not need to win a decade-long political battle to centralise the nation’s health data before research can begin. This is the architectural stance behind sovereign AI in healthcare: national capability on national data, without surrendering custody to anyone — including the platform vendor.
The decision framework: what you must own, what you should procure
Non-delegable: the governance layer
Four functions belong to the state or its statutory delegate regardless of platform choice: the legal framework and permitted purposes for data use; the access body that assesses applications (the Five Safes framework from the UK Office for National Statistics remains the reference model — see our explainer on the Five Safes); the accreditation of researchers and projects; and public transparency and opt-out administration. No vendor can supply legitimacy.
Procurable: the platform layer
The platform layer — secure workspaces, federated query and orchestration, harmonisation tooling, airlock and disclosure control, audit logging, identity integration — is where build-from-scratch programmes stall. This is deep, specialised engineering with a punishing maintenance tail: vocabulary updates, security patching, workload evolution from statistics to AI training. It is also where a decade of sector learning is already embodied in existing platforms. The honest comparison is therefore not “build versus buy” but “build, buy, or buy-and-deploy-sovereignly” — the third option being procured software running entirely within government-controlled infrastructure.
The hybrid rule of thumb
Own the governance, procure the platform, and insist the platform deploys federated within your boundary — with the contract written so that accreditation, audit rights, and disclosure-control rules remain the state’s to set and change. Programmes that build everything spend their first three years on infrastructure instead of research; programmes that outsource everything — including custody, on a vendor’s SaaS — discover after the UK Biobank incident what policy-only control is worth.
Build versus buy versus sovereign-deploy: the comparison
| Dimension | Build in-house | Buy (vendor SaaS, centralised) | Buy and deploy federated in national infrastructure |
|---|---|---|---|
| Time to first researcher | 3–5 years | Months | Months |
| Data custody | National — if not centralised into one repository | Vendor-hosted; custody effectively transferred | National throughout; data never leaves the source |
| Upfront cost profile | Large engineering programme before any research output | Low entry, subscription-based | Deployment plus subscription; no platform-engineering programme |
| Long-run maintenance | Permanent in-house team; key-person risk | Vendor-carried | Vendor-carried under national operational oversight |
| Egress control | Depends entirely on what you build | Policy-based; export workflows typically exist | Architectural — automated airlock, no record-level export path |
| EHDS Article 50 fit | Achievable, but you must engineer and evidence it | Depends on vendor hosting jurisdiction | Native — secure processing environment inside national jurisdiction |
| Scaling to new data custodians | Requires centralising each new source | Requires each custodian to surrender a copy | Federated — custodians join without moving data |
| Exit and lock-in risk | Locked into your own legacy | Data repatriation required at exit | Data already national; exit is a software swap |
What the reference programmes actually did
The instructive national examples are hybrids, not extremes. Genomics England — the UK’s national genomics programme — owns the governance, the participant relationship, and the data, and operates its research environment on Lifebit’s federated platform: researchers analyse in place, exports pass controlled review, and the data remains under Genomics England’s custody throughout. Singapore’s Ministry of Health takes the same posture for national health data, procuring federated platform capability while retaining sovereign control of the assets. Canada’s CanPath cohort federates across provincial holdings where centralisation would be constitutionally awkward. And the UK’s NHS Secure Data Environment programme illustrates the governance half: a national accreditation and policy framework under which platform capability is procured rather than hand-built. None of these programmes wrote their TRE from scratch; all of them kept custody.
The counter-examples are equally instructive, if less often named. Programmes that attempted full in-house builds have repeatedly discovered that the platform consumed the budget intended for the science: the Goldacre Review documented the UK’s history of duplicated, short-lived data infrastructure across the NHS — dozens of overlapping extract-and-disseminate arrangements, each rebuilt at cost, none accumulating into durable capability. The pattern generalises. A national TRE is not a project that ends; it is a service with a permanent security, harmonisation, and workload-evolution obligation, and the honest build-versus-buy comparison prices that obligation over a decade, not over the initial development contract. On that horizon, the total cost of ownership of an in-house build is dominated by the standing engineering team and the opportunity cost of delayed research output — both of which the hybrid model converts into a predictable service cost while the state’s own investment concentrates on governance, data quality, and the researcher community.
Common pitfalls and objections
Five recur in procurement. “Building keeps us in control” — control follows custody and operations, not authorship; an in-house platform on centralised architecture is less sovereign than procured software running federated in your cloud. “SaaS is faster” — faster to start, but the UK Biobank incident is the standing rebuttal on custody, and repatriating a national dataset at contract exit is a programme in itself. “We’ll centralise first, federate later” — centralisation is the step that takes years and burns custodian trust; federation removes it from the critical path entirely, as each custodian connects a Trusted Research Environment node without surrendering data. “Open-source components make build cheap” — components are not a platform; the integration, security hardening, and decade of maintenance are the cost, as the Goldacre Review’s analysis of duplicated NHS data infrastructure spending implies. “One national platform must mean one national database” — the opposite: a single research service over federated data is precisely what the architecture provides.
What to do next
Run the decision in this order. First, define the governance layer you must own — access body, permitted purposes, accreditation, transparency — and staff it, because no platform decision substitutes for it. Second, write custody into the procurement’s pass/fail criteria: data remains in national infrastructure; no record-level export path exists; every output passes automated disclosure control. Third, evaluate platforms against federation, not just security certifications — ask each vendor to demonstrate an analysis executing across two custodians without data movement. Fourth, sequence for early proof: one custodian, one cohort, one published study inside twelve months beats a five-year platform roadmap — early evidence recruits the next custodian more effectively than any mandate, and it gives ministers something concrete to defend when the programme’s budget is next reviewed. The EHDS timetable, the Goldacre Review, and the operating national programmes above give procurement teams both the mandate and the reference designs to move now.
Frequently asked questions
Should a government build or buy a national TRE?
The evidenced pattern is hybrid: own the governance layer (access body, permitted purposes, accreditation, transparency) and procure the platform layer, deployed federated inside government-controlled infrastructure. Building the platform from scratch typically costs three to five years before the first research output; vendor SaaS transfers custody. The hybrid delivers speed and sovereignty together.
How long does it take to stand up a national TRE?
With a procured federated platform deployed into national infrastructure, first-researcher access is achievable in months, with the schedule driven by governance readiness and data harmonisation rather than software. In-house builds historically run three to five years to equivalent capability.
Does buying a platform compromise data sovereignty?
Not if custody is the procurement criterion. Sovereignty is determined by where data resides, who operates the infrastructure, and what can leave — not by who wrote the software. A federated platform running inside government cloud keeps data in national jurisdiction, and data never leaves the source; a vendor-hosted centralised SaaS is the arrangement that compromises sovereignty.
What does the EHDS require of national TREs?
The European Health Data Space (Regulation (EU) 2025/327) requires secondary use of electronic health data to take place in secure processing environments — controlled settings that log activity and prevent extraction of record-level data (Article 50) — with health data access bodies in each member state issuing permits. A federated TRE inside national infrastructure satisfies the requirement by design.
Why choose a federated architecture for a national TRE?
Because it removes the two hardest problems: it eliminates the multi-year political and legal effort of centralising the nation’s health data, since custodians connect without moving records; and it makes bulk egress architecturally impossible rather than policy-forbidden — the failure mode the May 2026 UK Biobank incident exposed in centralised platforms.
What should the pass/fail criteria in a TRE procurement include?
At minimum: data remains within national jurisdiction and infrastructure; no record-level export workflow exists; every output passes automated disclosure control; the platform demonstrates federated analysis across two custodians without data movement; full audit logging; and a credible exit path in which the state already holds the data.
What did the Goldacre Review recommend about TREs?
The 2022 Goldacre Review (“Better, broader, safer: using health data for research and analysis”) recommended the UK shift from disseminating extracts of health data to providing access through a small number of accredited Trusted Research Environments — a recommendation NHS England subsequently implemented as its Secure Data Environment policy.
