Health Research Metadata Standards: DCAT & HealthDCAT-AP

Health research metadata standards are shared vocabularies for describing datasets so that researchers can find them, judge whether they fit a study, and request access without seeing a single record. The core European stack is DCAT (the W3C Data Catalog Vocabulary), DCAT-AP (its application profile for public sector data portals in Europe), and HealthDCAT-AP (the health extension built for the European Health Data Space). Together they turn a data catalog from a list of names into machine-readable descriptions that other catalogs, search tools, and access bodies can read.
Why metadata standards matter now
For most of the last decade, dataset metadata in health research was a courtesy. A biobank published a web page, a registry wrote a PDF data dictionary, and a hospital research office kept a spreadsheet of what it held. Each description used its own fields and its own words for the same things. A researcher looking for adult cohorts with linked genomic and primary care data had to read dozens of pages and email a dozen custodians to answer a question that should take minutes.
The European Health Data Space (EHDS) regulation, published in the Official Journal of the European Union in 2025, changes metadata from a courtesy into an obligation. It requires data holders to describe the datasets they hold, requires health data access bodies to publish national dataset catalogs, and connects those national catalogs to an EU-level catalog. It also introduces a data quality and utility label so that users can compare datasets on a common basis. None of that works unless every catalog speaks the same metadata language, which is exactly the job HealthDCAT-AP was built to do. Our guide to EHDS and the secondary use of health data covers the wider regulation.
The same pressure exists outside Europe. National programs, research consortia, and pharma federated networks all face the same discovery problem, and the FAIR principles (Findable, Accessible, Interoperable, Reusable) put rich, standard metadata at the center of the “Findable” requirement.
The standards, layer by layer
DCAT: the base vocabulary
DCAT is a W3C Recommendation for describing datasets and data services in catalogs. It defines a small set of classes: a catalog, a dataset, a distribution (a specific form in which the dataset is available, such as a file or an API), and a data service. Each class carries properties like title, description, publisher, keywords, theme, temporal coverage, spatial coverage, and license. DCAT is expressed in RDF (Resource Description Framework), which means a description is a set of linked statements rather than a flat form, and any system that understands RDF can merge and query descriptions from many sources. The current version, DCAT 3, added better support for versioning and dataset series.
DCAT-AP: the European profile
An application profile narrows a general vocabulary for a specific community. DCAT-AP, maintained through the European Commission’s SEMIC (Semantic Interoperability Community) action, states which DCAT properties are mandatory, recommended, or optional for European public sector portals, and which controlled vocabularies to use for values such as themes, languages, and file types. It is the reason the official EU open data portal can harvest metadata from national portals across member states. Many countries publish national extensions of DCAT-AP on top of it.
HealthDCAT-AP: the health extension
General open data metadata does not say enough about health data. A researcher needs to know the population covered, the coding systems used, whether records are linked, what legal basis governs reuse, who the access body is, and how the data is available: as a downloadable file, which is rare for sensitive data, or only inside a secure processing environment. HealthDCAT-AP extends DCAT-AP with properties for these points. It was developed through the European joint action work that prepared the ground for EHDS and is intended as the common format for dataset descriptions in EHDS catalogs.
Other standards you will meet
DCAT is not the only game. Schema.org’s Dataset type is what general web search engines read. DataCite metadata supports persistent identifiers (DOIs) for datasets. MIABIS (Minimum Information About BIobank data Sharing), developed in the BBMRI-ERIC biobanking community, describes biobanks, sample collections, and studies. HDR UK publishes a dataset metadata schema used by its Health Data Research Gateway. A mature catalog often maps its internal model to several of these at once.
The Data Harmonization angle: describing data you have not moved
Metadata standards and data harmonization are two halves of the same problem. Metadata describes what a dataset contains. Harmonization makes the contents comparable across datasets, typically by mapping source records to a common data model such as the OMOP CDM (Observational Medical Outcomes Partnership Common Data Model) or to FHIR (Fast Healthcare Interoperability Resources) resources. A catalog entry that says “diagnoses coded in ICD-10, mapped to OMOP CDM v5.4, with SNOMED CT standard concepts” is far more useful to a researcher than one that says “clinical data available on request.”
This is where a federated architecture changes what a catalog can say. In a centralized model, the catalog describes a copy, and the description ages from the moment the copy is taken. In a federated Trusted Research Environment (TRE), data never leaves the source, so the harmonized data sits at the custodian and the catalog can be generated from it directly. Row counts, date ranges, concept coverage, and completeness figures can be computed where the data lives and published as aggregate metadata, with small counts suppressed, without extracting a record. Lifebit’s approach is to treat harmonization to OMOP and FHIR as the step that makes that possible: once data is mapped to a common model, the same metadata queries run identically at every node.
Our practical guide to FAIR data principles explains why this matters for reuse, and the Lifebit Federated Trusted Research Environment page describes the platform that keeps data at source.
Comparing the main metadata standards
| Standard | Maintained by | Scope | Best for | Health-specific fields |
|---|---|---|---|---|
| DCAT | W3C | Any dataset or data service in a catalog | Base vocabulary for catalog interoperability | None |
| DCAT-AP | European Commission (SEMIC) | European public sector data portals | Harvesting between national and EU portals | None beyond themes |
| HealthDCAT-AP | European health data community, for EHDS | Health datasets under EHDS | EHDS national and EU dataset catalogs | Population, coding systems, access body, legal basis, quality information |
| Schema.org Dataset | Schema.org community | Datasets on the public web | Discovery through general search engines | Minimal |
| DataCite | DataCite | Citable research outputs | Persistent identifiers and citation | Minimal |
| MIABIS | BBMRI-ERIC community | Biobanks, sample collections, studies | Biobank directories and sample discovery | Sample types, disease focus, collection design |
A practical framework for building a standards-based catalog
- Inventory before you describe. List every dataset, its custodian, its legal basis for reuse, and how it can be accessed. Most organizations discover datasets they had forgotten and duplicates they had not noticed.
- Pick a primary profile and map outward. If you operate in or near the EU, adopt HealthDCAT-AP as your primary model now rather than retrofitting it later. Map from it to Schema.org for web discovery and to DataCite if you mint DOIs.
- Use controlled vocabularies for values. A free-text “Coding system” field will contain “ICD10”, “ICD-10”, and “icd 10 cm” within a month. Use the vocabularies the profile names, and use standard terminologies for clinical concepts.
- Generate what you can, curate what you must. Counts, date ranges, and concept coverage should be computed from harmonized data at the source on a schedule. Human effort belongs in descriptions, provenance, and access conditions, which no query can produce.
- Describe access, not just content. State who decides on access, which secure environment the data is analyzed in, and what outputs are allowed to leave. For sensitive data the distribution is usually a TRE, not a file.
- Publish in a harvestable form. Expose RDF or JSON-LD at stable URLs so national and EU catalogs can harvest you without manual re-entry.
- Version everything. Datasets change. Record when the description was generated and which release of the data it refers to.
Real-world examples
Europe already runs on DCAT-AP. The official EU open data portal harvests national catalogs that publish in the profile, which is the working proof that a shared application profile lets many independent catalogs appear as one. EHDS extends that model to health, with national health data access bodies publishing dataset descriptions that feed an EU catalog.
In biobanking, the BBMRI-ERIC Directory uses MIABIS to describe biobanks and collections across member countries, so a researcher can search for sample types and disease areas across many institutions from one place. In the UK, HDR UK’s Health Data Research Gateway publishes metadata for datasets from many custodians against a common schema and routes access requests to the right data controller. In each case, the value comes from the shared schema, not from any single catalog.
Common pitfalls and objections
“Our metadata is good enough already”
Internal metadata is usually good for internal users. The test is whether a machine at another organization can read it without a phone call. If your catalog cannot be harvested, it is not interoperable, however detailed it is.
Treating the catalog as a one-off project
A catalog written once and left alone becomes wrong within a year. Tie metadata generation to the data refresh cycle so descriptions update when the data does.
Publishing too much detail
Aggregate metadata can still disclose information about individuals when counts are small or when many fine-grained breakdowns are published together. Apply the same statistical disclosure control to published metadata that you apply to research outputs, including minimum cell sizes.
Confusing metadata with harmonization
A beautifully described dataset in a local coding scheme is still hard to use across sites. Metadata tells researchers what exists; harmonization lets them analyze it alongside other datasets. Plan both.
What to do next
Start with an honest inventory and a decision on your primary metadata profile. If EHDS applies to you, HealthDCAT-AP is the practical default. Then look at where your metadata comes from: if it is typed by hand from documents, the next step is harmonizing data at the source so that coverage and quality metadata can be computed and refreshed automatically. That is also the step that prepares you for federated analysis, where approved researchers run their queries in a secure environment at the custodian and only reviewed, aggregate results leave. For background on the environment itself, see what a Trusted Research Environment is.
Frequently asked questions
What is DCAT in health research?
DCAT, the Data Catalog Vocabulary, is a W3C Recommendation for describing datasets and data services in catalogs. In health research it provides the base structure, such as title, publisher, license, and distribution, on which health-specific profiles like HealthDCAT-AP add fields for population, coding systems, and access conditions.
What is the difference between DCAT-AP and HealthDCAT-AP?
DCAT-AP is the European application profile of DCAT for public sector data portals. HealthDCAT-AP extends DCAT-AP with properties needed to describe health datasets, including the population covered, the terminologies used, the health data access body, the legal basis for reuse, and quality information relevant to the European Health Data Space.
Is HealthDCAT-AP required under EHDS?
EHDS requires data holders to provide dataset descriptions and requires health data access bodies to publish dataset catalogs connected to an EU catalog. HealthDCAT-AP is being developed as the common format for those descriptions, so organizations in scope should plan to describe datasets in it.
Can a data catalog expose sensitive information?
Yes, if it publishes small counts or many detailed breakdowns. Catalogs should apply statistical disclosure control, such as minimum cell sizes and suppression, to aggregate metadata in the same way a TRE applies it to research outputs.
How do metadata standards relate to OMOP and FHIR?
Metadata standards describe a dataset from the outside, while OMOP and FHIR define how the data inside is structured and coded. A catalog entry can state that a dataset is mapped to OMOP CDM v5.4, which tells researchers the data is analysis-ready and comparable with other OMOP datasets.
Can metadata be generated in a federated TRE without moving data?
Yes. When data is harmonized at the source, aggregate metadata such as record counts, date ranges, and concept coverage can be computed where the data lives and published with small counts suppressed, so the catalog stays current without extracting individual records.
