Health Data Quality: Measure and Improve Real-World Data

Health data quality is measured by testing real-world data against defined dimensions, most commonly conformance, completeness, and plausibility, and against whether it is fit for a specific research question. Quality improves when those checks run automatically and repeatedly at the data source, results are published alongside the data, and fixes are made in the harmonization pipeline rather than by each analyst on their own copy.
Why data quality matters now
Real-world data (RWD), meaning data collected during routine care such as electronic health records (EHRs), claims, registries, and lab systems, is now used for decisions that used to rely only on trials. Regulators have responded by spelling out what they expect. The US Food and Drug Administration has published guidance on assessing EHR and medical claims data to support regulatory decision-making, with emphasis on relevance and reliability. The European Medicines Agency published a Data Quality Framework for EU medicines regulation, which describes dimensions including reliability, extensiveness, coherence, timeliness, and relevance. The European Health Data Space (EHDS) regulation introduces a data quality and utility label for datasets made available for secondary use.
The common message is that “we have a lot of data” is no longer an answer. Sponsors, regulators, and data access committees want evidence that a dataset is fit for the question being asked, and they want it documented before results arrive, not after.
What “quality” actually means for health data
The harmonized data quality framework
The most widely used vocabulary comes from a 2016 paper by Kahn and colleagues, which proposed a harmonized data quality terminology for EHR-based research. It groups checks into three categories.
- Conformance: does the data follow the expected format, structure, and value sets? For example, are dates valid dates and are diagnosis codes drawn from the expected vocabulary?
- Completeness: is data present where it should be? For example, what share of patients have a recorded sex, or a lab value for a test they were billed for?
- Plausibility: are the values believable? For example, no births after deaths, no pregnancies recorded in male patients, and heights within human range.
Each category can be assessed through verification, which checks data against internal rules and expectations, and validation, which compares data with an external reference such as a registry or published prevalence.
Fitness for purpose
A dataset is never simply “high quality”. It is fit or unfit for a purpose. A primary care dataset may capture diagnoses well and hospital procedures poorly. A claims dataset may record prescriptions filled but not whether they were taken. Quality assessment should always end with the question: for this study, which variables matter, and are they reliable enough?
The data harmonization angle
Quality and harmonization are tightly linked. Harmonization means transforming data from different sources into a common structure and vocabulary so it can be analyzed together. Many quality problems are created or revealed during that step: local codes that do not map to a standard concept, units that differ between labs, or a field that one hospital uses for a different purpose than another.
Harmonizing to a common model such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model has a large practical benefit: it makes quality checks reusable. The Observational Health Data Sciences and Informatics (OHDSI) community maintains the open-source Data Quality Dashboard, which implements the Kahn framework as thousands of standardized checks against OMOP tables, and Achilles, which profiles data for characterization. Once a source is in OMOP, the same checks run the same way at every site, and results can be compared. Our OMOP implementation checklist describes how to set up that pipeline.
In a federated setting this matters even more. In a federated Trusted Research Environment (TRE), each custodian keeps its data and analysis travels to it. Data never leaves the source. That means quality checks also have to run where the data lives. Lifebit’s approach is to run harmonization and quality checks inside each node, publish the summary results to researchers and data access committees, and keep record-level issues visible only to the custodian who can fix them. AI-assisted mapping speeds up the harmonization step, but mapping still needs review and quality checks to confirm it is right; our article on AI-automated OMOP harmonization at scale explains how the two fit together.
Who owns data quality
Quality work fails most often because nobody owns it. A workable split has three roles. The data custodian owns the source extract and fixes problems in source systems or extraction logic. The harmonization team owns mappings and transformation rules, and the checks that confirm they are correct. Researchers own the fitness-for-purpose assessment for their study, and they report issues they find back to the other two rather than patching them locally. Write this split into the data access agreement or the TRE operating model, with a named contact for each role and an agreed turnaround for fixes. Without it, the same problem is discovered and quietly worked around by every new study team, and published results diverge for reasons nobody can later explain.
A practical framework to measure and improve data quality
The table below maps common quality dimensions to example checks, how they are measured, and where fixes should happen.
| Dimension | Example check | How to measure | Where to fix |
|---|---|---|---|
| Conformance | Diagnosis codes map to standard concepts | Share of source codes mapped; count of unmapped records | Vocabulary mapping in the harmonization pipeline |
| Completeness | Sex, birth year, and visit dates recorded | Percentage populated per field, per site, over time | Source extraction logic; sometimes the source system |
| Plausibility | No events after death; lab values within human range | Count and rate of rule violations | Transformation rules; flag or exclude at source |
| Timeliness | Data refreshed on schedule; lag from event to record | Latest event date vs extract date | Refresh schedule and ETL monitoring |
| Coherence across sites | Similar prevalence of common conditions at comparable sites | Cross-site comparison of key rates | Site-specific mapping or coding practices |
| Validity against external data | Cancer incidence close to registry figures | Ratio to external benchmark | Linkage or capture gaps investigated with the source |
Step 1: Profile before you harmonize
Characterize the raw source first: row counts, value distributions, missingness, date ranges. This baseline lets you tell whether a later problem came from the source or from your transformation.
Step 2: Harmonize to a common model
Map data to OMOP or another agreed model, and record every mapping decision. Unmapped codes are not a failure to hide; they are a quality metric to report.
Step 3: Run standardized checks on every refresh
Automate the Data Quality Dashboard or equivalent checks so they run each time data is updated. A one-off assessment goes stale as soon as the next extract arrives.
Step 4: Assess fitness for each study
Before a study starts, check the specific variables it needs: exposure, outcome, key covariates, and follow-up time. Publish a short fitness-for-purpose note with the protocol.
Step 5: Fix problems at the right layer
Correct mapping errors in the pipeline, not in an analyst’s script. If the problem lies in the source system, report it back to the custodian. Fixes made downstream help one study; fixes made upstream help every study.
Step 6: Publish quality results with the data
Make quality summaries visible to researchers and access committees in the data catalog. This is also how custodians prepare for EHDS-style quality labeling.
Real-world example
DARWIN EU, the European Medicines Agency’s coordination center for real-world evidence, runs studies across a network of data partners that convert their data to the OMOP Common Data Model, and it uses standardized quality checks as part of onboarding data sources. OHDSI network studies follow a similar pattern: each site runs the Data Quality Dashboard locally and shares results, not patient records, before joining a study. Both show that quality assessment works well in distributed networks when the model and the checks are shared. For how this applies to multi-site research in practice, see running multi-site observational studies without moving data.
Common pitfalls
Treating quality as a single score
An overall pass rate hides the checks that matter for your study. Report results by dimension and by variable.
Checking once, at onboarding
Source systems change. Coding practices shift, new lab systems go live, and extracts break. Quality monitoring should be continuous.
Fixing data in analysis code
When each analyst cleans their own copy, results become hard to reproduce and fixes are never shared. Central fixes in the harmonization pipeline are more reliable.
Confusing missing with absent
In routine care data, a missing diagnosis may mean the patient does not have the condition, or that it was recorded somewhere else. Validation against external sources is the only way to tell.
Assuming federation hides quality problems
Federation does not stop quality assessment. It moves it to the source, where the custodian can act on it, and shares the summaries.
What to do next
Start by listing the quality checks you currently run, when they run, and who sees the results. Compare that list with the framework above and with the dimensions regulators now reference. Most organizations find the biggest gains in automating checks on every refresh and publishing results with the data. A federated TRE built on harmonized OMOP data gives you a consistent place to do both across every source you hold.
Frequently asked questions
What are the main dimensions of health data quality?
The most common framework uses conformance, completeness, and plausibility, assessed through verification and validation. Regulatory frameworks add dimensions such as relevance, reliability, timeliness, and coherence across sources.
What is the OHDSI Data Quality Dashboard?
It is an open-source tool from the OHDSI community that runs thousands of standardized quality checks against data in the OMOP Common Data Model, organized using the Kahn framework, and presents the results in a dashboard.
What does fit for purpose mean for real-world data?
It means the data is reliable and relevant enough for a specific research question. A dataset can be fit for one study and unfit for another, depending on which variables the study needs and how well they are captured.
How does harmonization affect data quality?
Harmonization reveals and sometimes creates quality issues, such as unmapped codes or unit mismatches. Harmonizing to a common model also makes quality checks reusable, so every site is measured the same way.
Can data quality be assessed in a federated network?
Yes. Each site runs the same standardized checks on its own data and shares summary results. Record-level issues stay with the custodian, who is best placed to fix them.
How often should data quality checks run?
Every time the data is refreshed. Source systems, coding practices, and extract logic change over time, so a one-off assessment quickly becomes out of date.
