Lifebit logo
BlogTechnologyClinical Trial Feasibility Analysis With Real-World Data

Clinical Trial Feasibility Analysis With Real-World Data

Dynamic abstract image of swirling blue light trails on a dark background.
Photo by Mahdi Bafande on Pexels

Clinical trial feasibility analysis with real-world data (RWD) means testing a protocol’s eligibility criteria against actual patient records — electronic health records, registries, and claims — before the trial opens, to establish how many eligible patients genuinely exist, where they are, and how criteria changes move those numbers. Done through a federated query layer, a sponsor can count eligible patients across dozens of hospitals and biobanks in minutes, without any patient-level data leaving the institutions that hold it.

Why feasibility is the highest-stakes estimate in a trial

Enrolment failure remains the most common way clinical trials lose time and money: industry and academic analyses consistently find that a large proportion of trials miss their original enrolment timelines, and that a meaningful share of sites recruit few or no patients at all. The conventional feasibility process invites this outcome, because it is built on questionnaires: sponsors ask investigators how many eligible patients they see, and investigators estimate optimistically from memory. Real-world data replaces that estimate with a measurement.

Regulators have simultaneously made RWD a first-class asset. The United States Food and Drug Administration (FDA) published its Real-World Evidence Framework in 2018 under the 21st Century Cures Act and has since issued final guidance on assessing electronic health records and medical claims data for regulatory decision-making. In Europe, the European Medicines Agency (EMA) operates DARWIN EU — the Data Analysis and Real World Interrogation Network — which runs studies across a network of data partners harmonised to the Observational Medical Outcomes Partnership (OMOP) Common Data Model. When the same regulators who will review your submission are building federated RWD networks themselves, using RWD to design the trial is no longer an innovation claim; it is baseline diligence.

Enrolment diversity adds a further regulatory pull. Under the Food and Drug Omnibus Reform Act of 2022, the FDA requires sponsors of most pivotal trials to submit Diversity Action Plans setting enrolment goals by demographic group, and has issued draft guidance on their format and content. A diversity goal is a feasibility question in disguise: it cannot be answered credibly with investigator questionnaires, but it can be answered with a federated query that returns eligible-patient counts broken down by age, sex, and geography across a network of institutions. Sponsors that already run RWD-based feasibility can attach measured numbers to their plans; those that do not are committing to targets they have no way to test.

The healthcare AI infrastructure angle

Feasibility analysis is usually discussed as a data problem, but in practice it is a healthcare AI infrastructure problem. The records that answer feasibility questions sit inside hospitals, academic medical centres, and biobanks that cannot — legally or ethically — export patient-level data to a sponsor. Any feasibility approach that begins with “send us your data” fails at the first data-protection review. The infrastructure that works is federated: the protocol’s criteria travel to each institution as a query, execute locally against harmonised records, and return only aggregate counts. This is the same architectural pattern described in federation for health data, and it is the foundation on which Lifebit’s federated Trusted Research Environment (TRE) is built: compute moves to the data, and data never leaves the source.

Two further pieces of healthcare AI infrastructure make federated feasibility reliable rather than merely possible. The first is harmonisation. A feasibility query is only trustworthy if “type 2 diabetes with an HbA1c above 8%” resolves to the same concept sets at every site, which is why serious networks standardise on OMOP CDM v5.4 and its vocabularies — the approach explained in what data harmonisation is. The second is AI-assisted criteria translation: protocol eligibility criteria are written in free text, and modern platforms use language models to convert them into structured, reviewable queries against the common data model, with a human confirming the mapping before it runs. Healthcare AI infrastructure of this kind turns feasibility from a six-week survey exercise into an interactive design session.

A practical framework for RWD-based feasibility

Step 1 — Structure the eligibility criteria

Decompose the protocol’s inclusion and exclusion criteria into computable elements: diagnoses, laboratory thresholds, medication exposures, procedures, and demographic bounds. Map each to standard vocabularies (SNOMED CT, LOINC, RxNorm) within the OMOP model. Flag the criteria that cannot be computed from structured data — performance status is a common example — and treat them as an attrition factor rather than pretending they resolve to zero.

Step 2 — Count, then relax

Run the full criteria set as a federated count across the candidate network. Then relax criteria one at a time and re-run, producing a sensitivity table: how many patients does each criterion remove? A criterion that eliminates 40% of otherwise-eligible patients while adding little scientific value is a protocol amendment waiting to happen — cheaper to fix before first-patient-in than after six months of under-enrolment.

Step 3 — Localise and select sites

Break the aggregate count down by institution and geography. Real-world feasibility routinely inverts assumptions here: the flagship academic centre may hold fewer protocol-eligible patients than a regional hospital with the right referral pattern. Select and tier sites on measured counts, not reputation. Geographic breakdown also feeds the protocol’s practical design: visit schedules that assume patients live near an academic centre look different when the eligible population turns out to be dispersed across community settings, and decentralised-trial elements can be justified with numbers rather than intuition.

Step 4 — Model attrition honestly

An eligible patient in the data is not an enrolled participant. Apply explicit attrition assumptions — contactable, consenting, passing the uncomputable criteria — and document them, so that the enrolment projection is an auditable model rather than a hopeful multiplication.

Step 5 — Re-run during the trial

Feasibility is not a one-off gate. Re-running the federated count monthly against refreshed data shows whether the eligible population is materialising as projected, and gives early, quantified warning before an amendment or rescue strategy becomes urgent.

Questionnaire-based versus federated RWD feasibility

DimensionQuestionnaire-based feasibilityFederated RWD feasibility
Evidence baseInvestigator recall and estimatesQueries against actual patient records
Typical turnaroundWeeks to months per survey roundMinutes to hours per query iteration
Criteria sensitivity testingImpractical — each variant needs a new surveyInteractive — relax a criterion and re-count
Data movementNone, but also no dataNone — aggregate counts only; data never leaves the institution
Site selectionReputation- and relationship-drivenMeasured eligible-patient counts per site
Known failure modeSystematic over-estimation of eligible patientsUnder-counting where source data is poorly harmonised
Regulatory alignmentNot evidence-generatingConsistent with FDA RWE guidance and EMA DARWIN EU practice

What this looks like at network scale

The infrastructure pattern is proven in national research programmes. Genomics England operates federated research infrastructure with Lifebit in which approved users analyse genomic and clinical data inside a governed environment rather than receiving extracts; the Canadian Partnership for Tomorrow’s Health (CanPath) opens Canada’s largest population cohort to researchers on the same principle. For a sponsor, the significance is that the hard part — institutions agreeing to expose harmonised, queryable data under governance they control — is a solved problem when the architecture is federated. A feasibility network is the same trust model applied to a commercial question: each hospital keeps custody, approves the use, and answers with counts. The OMOP Common Data Model that underpins such networks is described in detail in Lifebit’s OMOP guide.

Common pitfalls and objections

Three pitfalls recur. First, trusting counts from unharmonised data: if one site codes heart failure in local terms that the query misses, the network under-counts and the sponsor deprioritises a genuinely strong site — data quality assessment must precede feasibility, not follow it. Second, ignoring the uncomputable criteria: a count that ignores performance status or informed-consent likelihood overstates enrolment just as surely as an optimistic investigator. Third, treating privacy as an afterthought: feasibility queries must return aggregates with small-cell suppression, under the same disclosure controls a TRE applies to research outputs — otherwise the feasibility tool becomes a re-identification risk in its own right. The standard objection — “our therapeutic area is too specialised for RWD” — is usually an argument for RWD rather than against it: the rarer the population, the more expensive a wrong enrolment assumption becomes, and the more valuable a measured count across many institutions is. Rare-disease sponsors, who face the sparsest populations of all, were among the earliest adopters of network-scale feasibility queries for exactly this reason.

What to do next

Before the next protocol reaches final draft, run a pilot: take one recent trial that under-enrolled, restate its criteria as structured queries, and test them against an accessible RWD network or a single partner institution’s OMOP-harmonised data. Compare the measured eligible population with the original feasibility estimate. That single retrospective exercise typically settles the internal debate, and it produces the artefact a data-governance committee needs to approve a prospective feasibility capability: evidence that the question can be answered with aggregate counts alone, on healthcare AI infrastructure where data never leaves the source.

Frequently asked questions

What is clinical trial feasibility analysis with real-world data?

It is the practice of testing a protocol’s eligibility criteria as structured queries against real patient records — EHRs, registries, claims — to measure how many eligible patients exist and where, replacing investigator estimates with counts before the trial opens.

How does federated feasibility protect patient privacy?

The query travels to each institution and runs locally; only aggregate counts with small-cell suppression return to the sponsor. Patient-level data never leaves the institution, so no data-sharing agreement for record transfer is required.

Why is OMOP important for feasibility analysis?

The OMOP Common Data Model gives every participating institution the same table structures and vocabularies, so a single query means the same thing everywhere. Without harmonisation, cross-site counts are not comparable and site selection decisions built on them are unreliable.

Do regulators accept real-world data in trial design?

Yes. The FDA’s Real-World Evidence Framework and subsequent guidance address RWD fitness-for-use, and the EMA’s DARWIN EU network runs real-world studies across OMOP-harmonised data partners. Using RWD for feasibility is design-stage diligence, not a regulatory novelty.

Can feasibility counts predict actual enrolment?

Not on their own. Counts establish the eligible population; enrolment projections require explicit attrition assumptions for contact, consent, and criteria that cannot be computed from structured data. The value is that each assumption becomes visible and auditable.

What role does AI play in feasibility analysis?

AI models translate free-text eligibility criteria into structured queries, suggest concept-set mappings, and flag criteria with high patient-elimination impact. A human reviews every mapping before execution — the AI accelerates the workflow; the harmonised federated infrastructure makes the answer trustworthy.


Federate & Discover Everything. Move Nothing.


United Kingdom

3rd Floor Suite, 207 Regent Street, London, England, W1B 3HH United Kingdom

USA
228 East 45th Street, Suite 9E, New York, NY 10017, United States

© 2026 Lifebit Biotech Inc. DBA Lifebit. All rights reserved.

By using this website, you understand the information being presented is provided for informational purposes only and agree to our Cookie Policy and Privacy Policy.