Lifebit logo
BlogUncategorizedHealth Data Quality: Measure and Improve Real-World Data

Health Data Quality: Measure and Improve Real-World Data

Dynamic abstract image of swirling blue light trails on a dark background.
Photo by Mahdi Bafande on Pexels

Health data quality is measured by testing real-world data against defined dimensions, most commonly conformance, completeness, and plausibility, and against whether it is fit for a specific research question. Quality improves when those checks run automatically and repeatedly at the data source, results are published alongside the data, and fixes are made in the harmonization pipeline rather than by each analyst on their own copy.

Why data quality matters now

Real-world data (RWD), meaning data collected during routine care such as electronic health records (EHRs), claims, registries, and lab systems, is now used for decisions that used to rely only on trials. Regulators have responded by spelling out what they expect. The US Food and Drug Administration has published guidance on assessing EHR and medical claims data to support regulatory decision-making, with emphasis on relevance and reliability. The European Medicines Agency published a Data Quality Framework for EU medicines regulation, which describes dimensions including reliability, extensiveness, coherence, timeliness, and relevance. The European Health Data Space (EHDS) regulation introduces a data quality and utility label for datasets made available for secondary use.

The common message is that “we have a lot of data” is no longer an answer. Sponsors, regulators, and data access committees want evidence that a dataset is fit for the question being asked, and they want it documented before results arrive, not after.

What “quality” actually means for health data

The harmonized data quality framework

The most widely used vocabulary comes from a 2016 paper by Kahn and colleagues, which proposed a harmonized data quality terminology for EHR-based research. It groups checks into three categories.

  • Conformance: does the data follow the expected format, structure, and value sets? For example, are dates valid dates and are diagnosis codes drawn from the expected vocabulary?
  • Completeness: is data present where it should be? For example, what share of patients have a recorded sex, or a lab value for a test they were billed for?
  • Plausibility: are the values believable? For example, no births after deaths, no pregnancies recorded in male patients, and heights within human range.

Each category can be assessed through verification, which checks data against internal rules and expectations, and validation, which compares data with an external reference such as a registry or published prevalence.

Fitness for purpose

A dataset is never simply “high quality”. It is fit or unfit for a purpose. A primary care dataset may capture diagnoses well and hospital procedures poorly. A claims dataset may record prescriptions filled but not whether they were taken. Quality assessment should always end with the question: for this study, which variables matter, and are they reliable enough?

The data harmonization angle

Quality and harmonization are tightly linked. Harmonization means transforming data from different sources into a common structure and vocabulary so it can be analyzed together. Many quality problems are created or revealed during that step: local codes that do not map to a standard concept, units that differ between labs, or a field that one hospital uses for a different purpose than another.

Harmonizing to a common model such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model has a large practical benefit: it makes quality checks reusable. The Observational Health Data Sciences and Informatics (OHDSI) community maintains the open-source Data Quality Dashboard, which implements the Kahn framework as thousands of standardized checks against OMOP tables, and Achilles, which profiles data for characterization. Once a source is in OMOP, the same checks run the same way at every site, and results can be compared. Our OMOP implementation checklist describes how to set up that pipeline.

In a federated setting this matters even more. In a federated Trusted Research Environment (TRE), each custodian keeps its data and analysis travels to it. Data never leaves the source. That means quality checks also have to run where the data lives. Lifebit’s approach is to run harmonization and quality checks inside each node, publish the summary results to researchers and data access committees, and keep record-level issues visible only to the custodian who can fix them. AI-assisted mapping speeds up the harmonization step, but mapping still needs review and quality checks to confirm it is right; our article on AI-automated OMOP harmonization at scale explains how the two fit together.

Who owns data quality

Quality work fails most often because nobody owns it. A workable split has three roles. The data custodian owns the source extract and fixes problems in source systems or extraction logic. The harmonization team owns mappings and transformation rules, and the checks that confirm they are correct. Researchers own the fitness-for-purpose assessment for their study, and they report issues they find back to the other two rather than patching them locally. Write this split into the data access agreement or the TRE operating model, with a named contact for each role and an agreed turnaround for fixes. Without it, the same problem is discovered and quietly worked around by every new study team, and published results diverge for reasons nobody can later explain.

A practical framework to measure and improve data quality

The table below maps common quality dimensions to example checks, how they are measured, and where fixes should happen.

DimensionExample checkHow to measureWhere to fix
ConformanceDiagnosis codes map to standard conceptsShare of source codes mapped; count of unmapped recordsVocabulary mapping in the harmonization pipeline
CompletenessSex, birth year, and visit dates recordedPercentage populated per field, per site, over timeSource extraction logic; sometimes the source system
PlausibilityNo events after death; lab values within human rangeCount and rate of rule violationsTransformation rules; flag or exclude at source
TimelinessData refreshed on schedule; lag from event to recordLatest event date vs extract dateRefresh schedule and ETL monitoring
Coherence across sitesSimilar prevalence of common conditions at comparable sitesCross-site comparison of key ratesSite-specific mapping or coding practices
Validity against external dataCancer incidence close to registry figuresRatio to external benchmarkLinkage or capture gaps investigated with the source

Step 1: Profile before you harmonize

Characterize the raw source first: row counts, value distributions, missingness, date ranges. This baseline lets you tell whether a later problem came from the source or from your transformation.

Step 2: Harmonize to a common model

Map data to OMOP or another agreed model, and record every mapping decision. Unmapped codes are not a failure to hide; they are a quality metric to report.

Step 3: Run standardized checks on every refresh

Automate the Data Quality Dashboard or equivalent checks so they run each time data is updated. A one-off assessment goes stale as soon as the next extract arrives.

Step 4: Assess fitness for each study

Before a study starts, check the specific variables it needs: exposure, outcome, key covariates, and follow-up time. Publish a short fitness-for-purpose note with the protocol.

Step 5: Fix problems at the right layer

Correct mapping errors in the pipeline, not in an analyst’s script. If the problem lies in the source system, report it back to the custodian. Fixes made downstream help one study; fixes made upstream help every study.

Step 6: Publish quality results with the data

Make quality summaries visible to researchers and access committees in the data catalog. This is also how custodians prepare for EHDS-style quality labeling.

Real-world example

DARWIN EU, the European Medicines Agency’s coordination center for real-world evidence, runs studies across a network of data partners that convert their data to the OMOP Common Data Model, and it uses standardized quality checks as part of onboarding data sources. OHDSI network studies follow a similar pattern: each site runs the Data Quality Dashboard locally and shares results, not patient records, before joining a study. Both show that quality assessment works well in distributed networks when the model and the checks are shared. For how this applies to multi-site research in practice, see running multi-site observational studies without moving data.

Common pitfalls

Treating quality as a single score

An overall pass rate hides the checks that matter for your study. Report results by dimension and by variable.

Checking once, at onboarding

Source systems change. Coding practices shift, new lab systems go live, and extracts break. Quality monitoring should be continuous.

Fixing data in analysis code

When each analyst cleans their own copy, results become hard to reproduce and fixes are never shared. Central fixes in the harmonization pipeline are more reliable.

Confusing missing with absent

In routine care data, a missing diagnosis may mean the patient does not have the condition, or that it was recorded somewhere else. Validation against external sources is the only way to tell.

Assuming federation hides quality problems

Federation does not stop quality assessment. It moves it to the source, where the custodian can act on it, and shares the summaries.

What to do next

Start by listing the quality checks you currently run, when they run, and who sees the results. Compare that list with the framework above and with the dimensions regulators now reference. Most organizations find the biggest gains in automating checks on every refresh and publishing results with the data. A federated TRE built on harmonized OMOP data gives you a consistent place to do both across every source you hold.

Frequently asked questions

What are the main dimensions of health data quality?

The most common framework uses conformance, completeness, and plausibility, assessed through verification and validation. Regulatory frameworks add dimensions such as relevance, reliability, timeliness, and coherence across sources.

What is the OHDSI Data Quality Dashboard?

It is an open-source tool from the OHDSI community that runs thousands of standardized quality checks against data in the OMOP Common Data Model, organized using the Kahn framework, and presents the results in a dashboard.

What does fit for purpose mean for real-world data?

It means the data is reliable and relevant enough for a specific research question. A dataset can be fit for one study and unfit for another, depending on which variables the study needs and how well they are captured.

How does harmonization affect data quality?

Harmonization reveals and sometimes creates quality issues, such as unmapped codes or unit mismatches. Harmonizing to a common model also makes quality checks reusable, so every site is measured the same way.

Can data quality be assessed in a federated network?

Yes. Each site runs the same standardized checks on its own data and shares summary results. Record-level issues stay with the custodian, who is best placed to fix them.

How often should data quality checks run?

Every time the data is refreshed. Source systems, coding practices, and extract logic change over time, so a one-off assessment quickly becomes out of date.