Lifebit logo
BlogTrusted Research EnvironmentWhy Genomics Data Sharing Is Still Broken in 2026

Why Genomics Data Sharing Is Still Broken in 2026

Quick answer. Genomics data sharing is still broken in 2026 because the entire system still assumes data must move to be analyzed. Data Transfer Agreements take 18–36 months, pipelines fragment across incompatible formats and Eroom’s Law keeps tightening its grip on R&D productivity. The architectural fix is federated access: data stays in place, analysis travels to it and cohort wait times collapse from 24 months to 8 weeks.

Abstract genomic data visualisation representing the scale and fragmentation of modern research datasets
Photo by Google DeepMind on Pexels

Data Transfer Agreements take 18 months, pipelines fragment and Eroom’s Law keeps tightening its grip. Here’s what’s actually broken — and the one architectural shift that fixes it.

If you’re a VP of R&D watching cohort access timelines stretch past 24 months, a senior bioinformatician drowning in manual health data harmonization jobs, or a biobank director fielding a growing queue of Data Transfer Agreement requests, this article is written for you. Let’s be precise about what’s broken and what it actually costs.

Key genomics data-sharing stats: 18–36 months for a cross-border Data Transfer Agreement, 2–6 months average wait for a single data access application, $2.6B average cost to bring one drug to market, 275M+ patient records locked in siloed environments

The data exists, but the access doesn’t

The raw material for medical breakthroughs has never been more abundant. The cruel irony is that access to this data has gotten harder, not easier, as the datasets have grown.

Here’s the mechanism. As sequencing datasets scale from thousands to millions of participants, the sensitivity of the data increases. A genome is a permanent biological identifier, shared with family members, potentially revelatory about ethnicity, predispositions and ancestry. Legal frameworks have responded accordingly. GDPR in Europe, HIPAA in the United States and a patchwork of national genomic data protection laws have made the “just share it” model legally untenable for most organizations.

The response from the research community has been a proliferation of data governance structures: Data Access Committees (DACs), Data Transfer Agreements (DTAs), Material Transfer Agreements (MTAs) and Institutional Data Use Agreements. Every one of these exists for legitimate reasons. But collectively, they’ve created a system where accessing even public genomic data can take months.

The DTA problem: governance theater vs governed access

Of all the friction points in genomics data sharing, Data Transfer Agreements deserve particular scrutiny, because they represent a fundamental confusion between governance and movement.

A DTA is designed to ensure that sensitive data is only used for authorized purposes by authorized parties, under appropriate security conditions.

The problem is the mechanism. A DTA requires data to physically move from one institution to another, which triggers every legal, regulatory and security review process both parties maintain. For a granular breakdown of why this takes so long, see our dedicated piece on why Data Transfer Agreements take 18–36 months and what it costs you.

The result is a process that can take anywhere from months to years, depending on:

  • How many jurisdictions are involved (cross-border equals exponentially harder)
  • Whether the receiving institution’s security posture satisfies the data custodian
  • Whether legal teams at both institutions have aligned on contract language
  • Whether there’s an IRB or ethics board requirement on either side
  • How many people need to sign the final agreement
Comparison of five data-sharing approaches — Centralized SaaS, DIY On-Premise, Open Data Repository, Traditional DTA and Federated TRE (Lifebit) — across data movement, time to access, security posture and usability

Researchers increasingly select datasets based on what’s accessible rather than what’s most scientifically appropriate. Studies are designed around data that can be obtained within grant cycles, not the data that would give the strongest answer. Rare disease research is disproportionately harmed, since no single cohort has sufficient patient numbers.

Eroom’s Law and the hidden cost of data friction

Eroom’s Law describes a grim trend: the inflation-adjusted cost of developing a new, FDA-approved drug has roughly doubled every nine years since the 1950s. Today, the average cost to bring a single drug to market is approximately $2.6 billion.

Eroom's Law chart showing new drugs approved per $1B R&D falling from an indexed 100 in 1950 to near zero in 2026, while average R&D spend per approval rises above $2.5B

While many factors drive Eroom’s Law, one that rarely receives adequate attention is data infrastructure. Every month spent waiting on a DTA is a month of wasted team time. Every bioinformatics engineer running manual harmonization scripts is not building drug discovery pipelines. The drag is real, measurable and largely invisible in standard R&D efficiency metrics.

Let’s say that a senior bioinformatician at a major pharma company earns approximately $150,000–$200,000 per year. If 40% of their time is spent on data wrangling, format reconciliation and pipeline plumbing that shouldn’t require human attention, that’s $60,000–$80,000 per year in wasted labor, per person. Scale that across a team of 20, and you’re looking at $1.2–$1.6 million annually, just in one department, just in wasted data labor.

Five specific things that are broken

The problems in genomics data sharing fall into five distinct categories. Each has a different pain profile depending on who you are in the system.

1. The “data must move” assumption

Most existing data sharing infrastructure was built on the assumption that sharing means copying. This model made sense in 2001. But not in 2026, when datasets are measured in petabytes and privacy regulations have become a legal liability rather than just a technical challenge.

When data moves, every hop introduces risk: re-identification risk, breach risk, compliance risk and the practical risk that the copy becomes stale the moment it leaves the source. A genomic dataset linked to EHR data at a hospital is a living resource; the snapshot a researcher downloads is immediately out of date.

2. Fragmented pipelines and format wars

Ask any senior bioinformatician how much of their week is genuine analysis versus data preparation, and the answer will depress you. The genomics field has produced an extraordinary number of data formats — FASTQ, BAM, CRAM, VCF, PLINK, BCF — and an equally fragmented set of clinical data models. OMOP, FHIR, UK Biobank formats and dozens of institution-specific schemas all need to be reconciled before a single cross-cohort analysis can run.

Data harmonization work that should be automated, auditable and reproducible is instead done manually, inconsistently and often undocumented. When the next researcher arrives, they start from scratch.

3. Black-box AI that researchers can’t trust or audit

Genomics AI work is dominated by black-box models: tools that produce predictions without traceable provenance, trained on data whose composition is unclear, running in pipelines that can’t be audited for regulatory submission. This is a critical problem for anyone working toward clinical translation or regulatory approval.

A bioinformatician who can’t trace a variant call to the pipeline version that produced it, or a VP of R&D who can’t explain to regulators how a biomarker was derived, is building on sand.

4. Access inequality: who gets to use the data

Large pharmaceutical companies with dedicated regulatory affairs teams, established legal departments and existing relationships with major biobanks navigate DTA processes relatively well. Academic researchers at smaller institutions, researchers in lower-income countries and scientists working on rare diseases with small patient populations face dramatically higher barriers.

The practical result: the institutions with the most resources to process data are not necessarily the institutions with the most scientifically important questions to ask. Published data confirms that high-income countries produce 322 times more genomic studies per unit of disease burden than low-income countries. Access infrastructure is, in part, producing this disparity.

5. The compliance-usability tradeoff that shouldn’t exist

Organizations that take genomic data security seriously often make their data effectively unusable for research. Environments locked down to meet HIPAA, GDPR, ISO 27001 and FedRAMP standards can take weeks to provision, require weeks of researcher onboarding and generate so much friction that researchers route around the secure environment and use less controlled methods instead. For the strategic framing on how to escape this trade-off, see our guide to federated vs centralized data governance for genomics teams.

Side-by-side comparison of traditional sharing today vs the Lifebit federated access model — from 18-36 month DTAs, physical data copies, siloed datasets and $1M+ wasted bioinformatics labour to Governed Access Protocol in weeks, data never moves, 275M+ patient records as a single cohort and full pipeline auditability

The architecture that fixes it: data that never moves

The five problems above share a common root: they all assume, at some level, that the data must move. Remove that assumption, and the problem space changes entirely.

Federated analysis is not new as a concept. What is new is the maturity of production-grade infrastructure for implementing it at scale, across multiple institutions, with full compliance and auditability.

The technical argument is straightforward: why move a petabyte-scale genomic dataset across international borders when you can move a few kilobytes of analysis code instead?

The DTA Lifecycle: a timeline nobody wants — 5 sequential phases from research team submitting access request through DAC review, legal & security review, data extraction and onboarding, ending with data harmonisation at month 24+ before first analysis is possible

What this looks like in practice

Abstract architecture is easy to propose. What matters is whether it actually works at the scale and sensitivity of real genomic research. The evidence suggests it does.

A peer-reviewed study in The Lancet Oncology demonstrated something remarkable: a clinically meaningful discovery in which a 27% cure rate signal in breast cancer patients was made possible by federated analysis across Genomics England’s dataset, powered by the Lifebit Trusted Research Environment. The analysis ran where the data lived. The result was published in one of medicine’s most prestigious journals.

Boehringer Ingelheim reduced the time to validate gene-protein targets by 90% using federated access to large external biobanks. The NIH NLM Federated Data Workshop found that data transformation and harmonization that would “normally take years to achieve” could be completed through a single, centralized, auditable workflow. Flatiron Health achieved 10x faster global data connectivity for oncology data delivery.

The pattern is consistent: once you remove the assumption that data must move, the entire access timeline compresses — not by 10% or 20%, but by an order of magnitude.

The technical requirements behind trustworthy federation

For a federated model to genuinely replace a DTA — legally, ethically and operationally — it needs to satisfy the same requirements a DTA is designed to meet. This is where most early federated analysis attempts fell short: they solved the data movement problem without solving the governance problem, and data custodians rightly declined to rely on them.

A production-grade federated Trusted Research Environment requires:

1. For institutional buyers (biobank directors, program leads)

Full role-based access control with complete audit trails. Every researcher’s action logged, every data query recorded, every export policy-enforced and approved. Governance is built into the infrastructure, not documented in a PDF that nobody reads.

2. For technical users (bioinformaticians, data scientists)

Nextflow-native pipelines with full reproducibility. Automated OMOP and FHIR harmonization without manual mapping. Auditability on every pipeline execution — every variant call is traceable to the tool version, parameters and data version that produced it.

3. For economic buyers (VPs R&D, CSOs)

Certifiable compliance — FedRAMP, HIPAA, ISO 27001, SOC 2 — built in, not bolted on. Results-guaranteed deployment with contractual SLAs. Cohort access reduced from 24 months to 8 weeks. Bioinformatics time recovered and redirected to discovery.

4. The airlock problem

Zero data movement means nothing if there’s an unmonitored download button. A production federated environment requires AI-automated export governance — every output reviewed, every export policy-enforced, every approval logged. This is the layer most implementations miss.

What needs to change at the institutional level

Policy and institutional behavior also need to change, and in several areas the policy environment is finally catching up. The NIH proposed significant revisions to its Genomic Data Sharing Policy in late 2025, with provisions for updated access models that better reflect the realities of federated analysis. The EU Health Data Space framework, which came into force in 2024, explicitly accommodates federated access as a preferred model for secondary use of health data. These are positive signals.

Biobank directors and government program leads who want to reduce their DTA backlog and increase researcher throughput without increasing security risk have a clear path: deploy a federated TRE that converts incoming Data Transfer Agreement requests into Governed Access Protocols.

The researchers who can’t wait any longer

These are not hypothetical scenarios. They are the everyday reality of genomics research in 2026 — a field that has the data, the compute and the scientific questions to make extraordinary progress, held back by infrastructure that was designed for a different era.

The good news is that the architecture exists to fix it. The question now is not whether federated access works — because it does. The question is how quickly institutions can migrate from a data-movement paradigm that no longer serves them to one that does.

What to do next

If you’re a VP of R&D or Chief Scientific Officer, the next time your team tells you cohort access will take 18 months, ask whether that timeline is genuinely necessary or an artifact of data infrastructure decisions made a decade ago. Federated TREs are proven to compress access timelines by an order of magnitude. The ROI case is not complicated.

See the federated TRE in action

Learn how Lifebit turns Data Transfer Agreements into Governed Access Protocols — reducing cohort access from 24 months to 8 weeks, with zero data movement and full compliance.

Book a demo →

FAQs about genomics data sharing

If data never moves in a federated architecture, how can researchers combine and analyze datasets stored across entirely different global biobanks?

This is the core strength of a Federated Data Platform powered by a Trusted Research Environment (TRE). Instead of physically pulling multiple datasets into a single local server, the platform acts as a unified orchestration layer. Researchers write or select their analytical pipelines once. The platform then securely distributes this code to execute locally within each data custodian’s isolated cloud infrastructure. The heavy compute happens right where the data lives. Once the local analyses are complete, only the highly aggregated, non-identifiable results are sent back to the researcher’s central dashboard to be combined.

How does the “broken” data sharing system create scientific bias and stall rare disease research?

When Data Transfer Agreements (DTAs) take up to two years, researchers face a brutal choice: stall their project or change their hypothesis. To hit grant deadlines or product launch windows, scientists increasingly select datasets based on accessibility rather than scientific relevance. They design studies around suboptimal, easily obtainable data. This catastrophic friction disproportionately harms rare disease and oncology research. Because rare disease patient populations are globally scattered, no single biobank holds a statistically significant cohort. Discovering meaningful genetic signals requires aggregating data across multiple international sites. Under the old “copy-and-paste” model, negotiating five separate international DTAs creates a compounding legal gridlock that effectively kills the study before it begins.