What Are Computable Phenotypes? Sharing Across Sites

A computable phenotype is a precise, machine-executable definition of a clinical condition or characteristic, such as “adults with type 2 diabetes”, built entirely from coded data in electronic health records and related sources, without needing a clinician to review each chart. Computable phenotypes are shared across sites by expressing them against a common data model and standard vocabularies, publishing them in a phenotype library, and running the same definition locally at each site so that only counts or results, not patient records, travel back.
Why computable phenotypes matter now
Almost every study that uses real-world data starts by finding the right patients. Who has the disease? Who received the drug? Who had the outcome? If each site answers those questions differently, a multi-site study compares different populations under the same label, and its results cannot be trusted or reproduced.
That problem has grown as research networks have grown. Large distributed programs now run the same study across dozens of hospitals and countries. Regulators reviewing real-world evidence ask how exposures and outcomes were defined and validated. The European Health Data Space (EHDS) will make secondary-use data from many member states available for research, which only works if definitions travel with the question. A shared, tested computable phenotype is how a research team makes “patients with heart failure” mean the same thing in every dataset it touches.
What a computable phenotype contains
The building blocks
A typical rule-based phenotype combines several elements.
- Code lists (value sets): sets of codes that represent a concept, drawn from vocabularies such as ICD-10 for diagnoses, SNOMED CT for clinical terms, LOINC for lab tests, and RxNorm for medications.
- Logic: how those codes combine, for example “two diagnosis codes at least 30 days apart, or one diagnosis plus a prescription for a relevant drug”.
- Thresholds: lab values or measurements, such as an HbA1c result above a stated level.
- Time windows: when events must occur relative to each other or to an index date.
- Exclusions: conditions or events that rule a patient out, such as a type 1 diabetes code in a type 2 diabetes phenotype.
Rule-based and machine learning phenotypes
Most shared phenotypes are rule-based, which makes them transparent and portable. Some groups also build probabilistic or machine learning phenotypes that score patients on their likelihood of having a condition, often using many features. These can be more accurate but are harder to share, because the model depends on the data it was trained on.
Validation
A phenotype is only as good as its measured performance. Validation usually compares phenotype results with a reference standard, most often expert chart review on a sample, and reports positive predictive value (the share of identified patients who truly have the condition) and, where possible, sensitivity (the share of true cases the phenotype finds). Performance measured at one site does not guarantee the same performance elsewhere, because coding habits differ.
The data harmonization angle
Computable phenotypes are where harmonization pays off. A phenotype written against one hospital’s local codes and table structure cannot run anywhere else without being rewritten, and every rewrite introduces differences. A phenotype written against a common data model with standard vocabularies can run unchanged at every site that uses that model.
The Observational Medical Outcomes Partnership (OMOP) Common Data Model is the most widely used example. The Observational Health Data Sciences and Informatics (OHDSI) community builds cohort definitions in its ATLAS tool, stores them as structured, machine-readable definitions, and shares them through the OHDSI Phenotype Library. Because OMOP maps source codes to standard concepts and includes concept hierarchies, a definition can say “any descendant of this heart failure concept” and pick up the right local codes at each site. Our overview of the OMOP Common Data Model explains the vocabulary layer in more detail.
Other standards serve related needs. Clinical Quality Language (CQL), an HL7 standard, expresses logic for quality measures and decision support, often against FHIR data. In the US, the Value Set Authority Center publishes value sets used in quality measurement. Whichever standard is used, the principle is the same: harmonize the data first so the definition can be shared as-is.
Federation completes the picture. In a federated Trusted Research Environment (TRE), each site keeps its harmonized data and the phenotype is sent to run locally. Data never leaves the source. Researchers receive counts, characteristics, or study results, subject to disclosure checks, rather than patient-level extracts. Lifebit’s federated TRE combines AI-assisted OMOP harmonization at each node with federated execution, so a single cohort definition can be run across nodes and return aggregate results through an airlock. Our guide to AI-generated cohorts on federated data shows how natural-language questions can be turned into reviewable cohort definitions that run this way.
What researchers get back from each site
When a phenotype runs in a federated network, the useful output is more than a single count. Well-designed networks return a small, agreed set of aggregate results from each site: the number of patients meeting the definition, a breakdown by age band and sex, the codes that contributed most to inclusion, and counts over time. Small cells are suppressed according to the network’s disclosure rules. This lets the study team compare sites side by side and spot definitions that behave oddly. A site where one rarely used code accounts for most cases, or where counts jump in a single year, usually points to a coding change or a mapping problem rather than a real difference in disease. Catching that before the main analysis is far cheaper than explaining it in peer review.
Phenotypes as part of the study record
Treat the phenotype as a formal study artifact, alongside the protocol and analysis plan. Store the exact version that ran, the vocabulary release used at each site, and the diagnostics output. Reviewers, regulators, and later research teams can then see precisely who was studied and repeat the selection on updated data.
How phenotypes are shared: a comparison of approaches
| Approach | How the definition travels | Portability | Main risk |
|---|---|---|---|
| Narrative description in a paper | Text and a code list in a supplement | Low: every site reimplements it | Different interpretations produce different cohorts |
| Local SQL scripts | Code written for one database | Low to moderate: needs rewriting for other schemas | Hidden local assumptions and errors |
| Common data model definition (for example OMOP cohort JSON) | Structured definition executed by standard tools | High across sites using the same model | Depends on mapping quality at each site |
| Standards-based logic (for example CQL on FHIR) | Formal logic with published value sets | High where FHIR data and engines exist | Uneven FHIR coverage of research data |
| Federated execution in a TRE | Same definition dispatched to each site; results returned | High, and no data movement | Requires shared model and output rules |
A practical framework for building and sharing phenotypes
- Start from a library. Check the OHDSI Phenotype Library, the Phenotype KnowledgeBase (PheKB), or the HDR UK Phenotype Library before writing a new definition. Reusing a validated phenotype saves time and makes results comparable with prior work.
- Write against a common model. Define logic in standard concepts and structures rather than local codes.
- Document intent as well as logic. State what clinical idea the phenotype is meant to capture, what it deliberately excludes, and why.
- Run diagnostics at every site. Tools such as OHDSI CohortDiagnostics show which codes drive inclusion, how counts vary by site, and where definitions behave unexpectedly.
- Validate where you can. Chart review on a sample, or probabilistic methods such as OHDSI’s PheValuator, give estimates of accuracy.
- Version and publish. Give every change a new version, and record which version each study used.
Real-world examples
The eMERGE (Electronic Medical Records and Genomics) network developed many of the phenotypes on PheKB and showed that carefully built algorithms could be transferred across health systems with documented performance. OHDSI network studies routinely run shared cohort definitions across data partners in many countries, with each site executing locally and returning aggregate results. In the UK, the HDR UK Phenotype Library collects definitions used across national research. For trial feasibility and recruitment, the same approach lets sponsors check how many eligible patients exist across sites before committing; see federated TRE for clinical trial cohort discovery.
Common pitfalls
Assuming codes mean the same thing everywhere
A diagnosis code used for billing at one hospital may be used for rule-out tests at another. Site diagnostics catch this; assumptions do not.
Skipping validation
An unvalidated phenotype is a hypothesis. Report what is known about its accuracy, even if only at one site.
Overfitting to one dataset
Definitions tuned to one site’s data often lose accuracy elsewhere. Test on more than one source before publishing.
Losing track of versions
Small edits change cohorts. Without versioning, two studies can cite “the same” phenotype and use different logic.
What to do next
List the conditions, exposures, and outcomes your organization studies most often, and check whether each has a published, versioned definition that runs on your harmonized data. Where it does not, adopt one from a public library or build one against a common model and publish it. Harmonized data in a federated TRE is what lets those definitions run consistently across every site you work with.
Frequently asked questions
What is a computable phenotype?
It is a precise definition of a clinical condition or characteristic, written so a computer can apply it to coded health data without clinician review of each record. It usually combines code lists, logic, thresholds, time windows, and exclusions.
How is a computable phenotype different from a code list?
A code list is one ingredient. A phenotype adds the logic around it, such as how many codes are needed, over what time period, alongside which lab results or medications, and what excludes a patient.
Where can I find existing computable phenotypes?
Public libraries include the OHDSI Phenotype Library, the Phenotype KnowledgeBase (PheKB) developed with the eMERGE network, and the HDR UK Phenotype Library.
How are phenotypes validated?
Most often by comparing phenotype results with expert chart review on a sample of patients and reporting positive predictive value and sensitivity. Probabilistic tools can estimate accuracy when chart review is not possible.
Why does a common data model help share phenotypes?
A common data model gives every site the same structure and standard vocabularies, so one definition can run unchanged everywhere. Without it, each site must rewrite the logic for its own database.
Can phenotypes be run without moving patient data?
Yes. In a federated network, the same definition is sent to each site, runs locally on harmonized data, and returns only aggregate results after disclosure checks.
