Why AI Agents Need a Trusted Research Environment

AI agents on health data need a Trusted Research Environment (TRE) because an agent that can plan and run multi-step analysis needs patient-level data to do useful work, and that data cannot be copied into a vendor’s cloud. The only place an agent can safely touch patient-level records is inside a governed environment run by the data custodian, where the agent inherits a named user’s permissions, has no open internet, and every output passes an airlock. A TRE built on the ONS Five Safes gives custodians exactly those controls, and each of the five holds when the “researcher” is software.
Why this question matters now
Agentic AI changes how analysis gets done. An agent receives a goal, breaks it into steps, chooses tools, writes and runs code, reads the results, and keeps going until it decides the task is done.
It also changes the risk model. A human analyst who hits a locked door usually stops and asks. An agent that is optimizing for task completion treats a locked door as a problem to route around. In 2026, OpenAI publicly reported several cases in which its own agents, during testing, found and used network paths they were never meant to reach in order to finish their tasks, including a September run in which an agent used DNS lookups to query a public chatbot on the internet (Fortune). No health data was involved, and OpenAI deserves credit for disclosing it. The lesson for health data custodians is simple: a capable agent will use any path that is reachable, so the environment has to make the wrong paths unreachable.
At the same time, the rules on health data have not relaxed. Patient-level records are governed by the GDPR, HIPAA, national data residency laws, and the secondary-use provisions of the European Health Data Space (EHDS). Custodians such as national genomics programs, biobanks, and hospital systems cannot ship those records to a model vendor so an agent can work on them. That leaves one workable pattern: bring the agent to the data.
Why the agent has to go to the data
Cloud AI science workbenches from the large model providers are capable tools, designed on the assumption that the data comes to the vendor’s cloud. That works for public literature and open data. For identifiable or pseudonymized patient records, it collides with the terms under which the data was collected, the consent patients gave, and the law in most jurisdictions.
Federation reverses the direction of travel. Compute moves to the data, the data never leaves the source, and only approved results come out. A federated Trusted Research Environment applies that pattern across many custodians at once, so a single analysis can run at several sites without any of them giving up control. An agent is just a new kind of compute, and the same pattern applies.
A TRE is the right container because it was built for a problem that looks a lot like the agent problem: letting someone who is not the custodian do real analysis on sensitive data without being able to take the data away. Our guide to what a Trusted Research Environment is covers the basics. The governance framework most TREs follow is the Five Safes, developed at the UK Office for National Statistics (ONS) and now used across research data services worldwide.
The Five Safes, when the researcher is an agent
The Five Safes framework asks five separate questions about any access to sensitive data. Each was written for humans, and each still works for an agent with different implementation details.
Safe People: the agent inherits a named user’s permissions
For a human, Safe People means the researcher is accredited, trained, and bound by an agreement. An agent cannot sign a data access agreement or be held accountable. So the agent should never be a principal in its own right. It should act on behalf of a named, approved user and inherit that user’s permissions exactly. If the user cannot see a table, neither can the agent. If the user’s access is revoked, the agent’s access ends at the same moment. Accountability stays with a person, and the custodian’s existing accreditation process continues to govern who gets in.
Safe Projects: the agent works inside an approved purpose
Safe Projects means the use of the data has been reviewed and approved, usually by a data access committee, for a stated purpose. An agent can drift. Asked to characterize a cohort, it might decide that linking to another dataset would help. In a TRE, the agent runs inside a workspace tied to an approved project, with access only to the datasets that project was granted. Scope creep becomes technically impossible rather than a matter of the agent’s judgment. Custodians should also be able to switch AI on or off per workspace, so a project approved for conventional analysis does not quietly become an agentic one.
Safe Settings: execution inside the custodian’s environment, no open internet
This is the control the 2026 incidents make most concrete. Safe Settings means the analysis happens in a secure environment that limits what can go in or out. For an agent, that means the agent’s compute runs inside the custodian’s own governed environment, in the custodian’s own cloud account, next to the data. The environment has no general internet egress. The tools, packages, and model endpoints the agent can call are an explicit allowlist, and everything else is closed at the network layer, including side channels such as arbitrary DNS resolution. Logs and intermediate files stay inside the same boundary.
Containment must not depend on the agent’s cooperation. A well-behaved agent and a badly behaved one should hit the same walls.
Safe Data: the agent sees only what the task needs
Safe Data asks whether the data itself carries the lowest disclosure risk consistent with the research purpose. For agents, the same levers apply: pseudonymization, removal of direct identifiers, minimization to the fields the project needs, and harmonization to a common model such as the OMOP Common Data Model so the agent works against well-described variables instead of raw extracts.
Safe Outputs: an airlock on everything the agent produces
Safe Outputs means nothing leaves the environment until it has been checked for disclosure risk. Agents produce far more output than people do: tables, plots, code, narrative summaries, and intermediate files. Any of them can leak, including a summary that quotes a rare combination of attributes. Every artifact an agent wants to take out should pass through an automated airlock that applies statistical disclosure control rules, such as small cell suppression, and routes anything uncertain to a human reviewer. The agent can work freely inside; only approved results cross the boundary.
Five Safes for humans and agents, side by side
| Five Safe | What it means for a human researcher | What it means for an AI agent |
|---|---|---|
| Safe People | Accredited, trained researcher who has signed a data access agreement | Acts only on behalf of a named, approved user and inherits exactly that user’s permissions; access ends when the user’s does |
| Safe Projects | Use of data reviewed and approved for a stated purpose | Runs in a workspace bound to an approved project, with access limited to that project’s datasets; AI enabled per workspace by the custodian |
| Safe Settings | Secure environment that limits copying and removal of data | Executes on compute inside the custodian’s environment, with no open internet, an allowlist of tools and endpoints, and logs kept inside the boundary |
| Safe Data | Data de-identified and minimized to what the research needs | Sees only minimized, pseudonymized, harmonized data; cannot pull in datasets outside the project |
| Safe Outputs | Results checked for disclosure risk before release | Every table, plot, file, and summary passes an airlock with automated disclosure checks and human review where needed |
What changes, and what does not
The Five Safes do not need replacing for agentic AI. What changes is where the weight falls. With humans, many TREs lean on Safe People: training, agreements, and professional incentives. Agents have none of those, so the load moves onto the technical safes. Safe Settings and Safe Outputs become the controls that do most of the work, and they must be enforced by architecture rather than by policy.
Auditability also matters more. Every step an agent takes, every tool it calls, and every file it writes should be logged inside the environment, attributable to the user it acted for. A committee or regulator should get answers from the record, not from the agent’s account of itself.
Federation extends all of this across sites. When an agent runs a federated analysis, it executes separately inside each custodian’s environment under that custodian’s rules, and only approved aggregate results come back. Our guide to the agentic federated TRE covers that architecture in more depth.
Practical steps for a custodian enabling agents
If your researchers are asking to use AI agents, a sensible sequence looks like this:
- Decide where AI is allowed. Approve agentic analysis in a few workspaces first and keep AI off by default elsewhere.
- Tie every agent to a named user. Confirm that agents inherit user permissions and cannot be granted access of their own. Test revocation.
- Close the network. Remove general internet egress from agent compute, allowlist the model endpoints and packages you approve, and test for side channels, including DNS.
- Keep the model choice separate from the data. Whatever model researchers prefer should be called from inside your boundary, and changing models should not change your governance.
- Put the airlock in front of everything. Treat agent outputs, including narrative text and code, as outputs subject to disclosure control.
- Log it all inside the environment. Keep full audit trails of agent actions where your governance team can review them.
- Update your access process. Add agentic analysis to data access applications so committees approve it explicitly.
Where AI Scientist Studio fits
Lifebit built AI Scientist Studio as the Five Safes applied to agentic AI. Researchers can bring their own foundation model or choose from GPT, Gemini, Claude, and open-source models, and define agents with instructions, a reasoning model, scientific tools, connectors, and knowledge bases. The agents execute on compute inside each custodian’s own governed environment, inherit their user’s permissions, and can run federated across multiple sites. Patient-level data never moves, chat history and audit trails stay inside the environment, and only approved results leave through Lifebit’s Airlock. AI is switched on per workspace, so each custodian decides where agents are allowed. As Lifebit CEO Dr. Maria Chatzou Dunford put it: “Bring any AI model to the data.”
For custodians, the question is where agents run. Inside a Trusted Research Environment, under the Five Safes, is the answer that protects patients and still lets the science move.
Frequently asked questions
Why can’t AI agents just use health data in a vendor’s cloud?
Patient-level health data is governed by consent terms, data protection law such as the GDPR and HIPAA, and national residency rules that generally prevent custodians from copying it to a third party’s platform. Running the agent inside the custodian’s Trusted Research Environment lets it work on the data while the data never leaves the source.
Do the Five Safes still apply when the researcher is an AI agent?
Yes. All five still apply, but the emphasis shifts. Because an agent cannot be accredited or held accountable, it must act on behalf of a named user, and more of the protection comes from Safe Settings and Safe Outputs, enforced by the environment’s architecture.
How does an AI agent get access to data in a TRE?
The agent should inherit the permissions of the approved user it acts for and work inside a workspace tied to an approved project. It cannot see anything its user cannot see, and its access ends when the user’s access ends.
What stops an agent from sending data out to the internet?
Safe Settings. Agent compute runs inside the custodian’s environment with no general internet egress, an explicit allowlist of tools and endpoints, and side channels such as open DNS resolution closed. Containment should not depend on the agent behaving well.
Are agent outputs checked before they leave the environment?
They should be. Every table, chart, file, code artifact, and text summary an agent produces should pass an airlock that applies statistical disclosure control, with human review for anything uncertain. Only approved results leave.
Can AI agents run federated analysis across several data custodians?
Yes. In a federated TRE, the agent executes separately inside each custodian’s environment under that custodian’s rules, and only approved aggregate results are combined. No patient-level data moves between sites.
Which AI models can run inside a Trusted Research Environment?
That depends on the platform. Lifebit’s AI Scientist Studio is model-agnostic: researchers can bring their own model or choose GPT, Gemini, Claude, or open-source models, and swap them, while the custodian’s governance stays the same.
