Lifebit logo
BlogIndustryThe NIH S-Index: Measuring Data Sharing That Matters

The NIH S-Index: Measuring Data Sharing That Matters

Abstract close-up of neon blue light in dark setting, highlighting modern artistic design.
Photo by Francesco Ungaro on Pexels

The S-index is a proposed metric — the subject of a $1 million NIH prize competition — that would credit researchers for high-quality, reusable data sharing in the way the h-index credits them for publications and citations. Rather than counting deposits, an S-index scores whether shared data is findable, well-annotated, standards-aligned, and actually reused by others, turning data sharing from a compliance obligation into a recognised scholarly contribution.

Why NIH is putting prize money behind a data-sharing metric

The NIH Data Management and Sharing Policy (in effect since January 2023) obliges funded researchers to plan and perform data sharing. It created real momentum — and exposed a gap. Policy can require that data be deposited; it cannot make researchers want to share well. Academic careers run on attribution, and today a widely reused dataset earns its creator a fraction of the recognition a moderately cited paper does. Datasets are rarely cited with the discipline that publications are, and the contribution of the people who prepare, document, and steward data is close to invisible in promotion decisions.

The Data Sharing Index challenge — a two-phase, open-innovation competition led from within NIH, with Phase 1 winners announced in July 2026 and a larger implementation phase to follow — is an attempt to close that gap with incentives rather than mandates. The premise is simple: measure the quality and downstream impact of shared data, attach the measurement to the researcher, and the behaviour follows. The design brief points at the FAIR principles (Findable, Accessible, Interoperable, Reusable), timeliness of release, annotation quality, and evidence of downstream reuse — publications, further datasets, even patents — as the raw material for the score.

H-index vs S-index — what changes

DimensionH-index (publications)S-index (data sharing)
What is countedPapers and the citations they accumulateShared datasets and the reuse they enable
Attribution mechanismMature: citation is a settled scholarly normImmature: data citation is inconsistent and often absent
Quality signalPeer review before publicationMetadata quality, standards alignment, documentation, provenance
Known failure modesSalami-slicing, citation cartelsDeposit-and-forget, metric gaming, overcounting derivative reuse
Who gets creditListed authorsOpen question — ideally including data stewards and curators, not only the PI

The right-hand column’s open questions are the interesting part. A metric that only counted deposits would reward volume over usefulness; a metric that counted reuse must first solve how reuse is observed — which is harder than it sounds, and hardest of all for sensitive data.

The federated blind spot in reuse metrics

Here is the structural problem an S-index has to solve to be fair: the easiest reuse to measure is reuse in a centralised, open repository, where downloads and citations can be logged in one place. But a large and growing share of the world’s most valuable health data cannot be centralised or openly downloaded at all. Genomic cohorts, national health records, and clinical datasets are analysed inside governed environments — Trusted Research Environments and federated networks where data never leaves the source and only vetted aggregate results come out.

If a data-sharing metric implicitly equates “reusable” with “downloadable”, it will systematically undervalue exactly the custodians who share the most sensitive, most valuable data in the most responsible way. A biobank whose dataset powers a hundred federated analyses — each run at source, each output reviewed through an automated airlock — has enabled more science than many open datasets ever will, yet generates no download logs for a naive metric to count.

What federated infrastructure can contribute

The irony is that governed environments are better positioned to evidence reuse than open ones — because everything that happens inside them is already recorded. A federated Trusted Research Environment maintains, as a matter of governance, precisely the records an S-index needs:

  • Provenance: which dataset versions were queried, by which approved project, under which data-access agreement.
  • Auditable reuse events: every federated analysis is logged at the source — a verifiable, non-gameable record that reuse occurred, without exposing a single record.
  • Standards alignment: data harmonised to common models such as OMOP CDM is measurably more interoperable, and the harmonisation itself is documented work that a fair metric should credit.
  • Output lineage: airlock-reviewed results connect downstream publications back to the governed data assets that produced them.

In other words, the question “how do you measure data reuse without moving the data?” has the same answer as “how do you analyse data without moving it?” — you compute the evidence where the data lives, and export only the aggregate. An S-index calculation can be treated as one more federated query. That framing also addresses the attribution question: when value is created across multiple custodians in a single federated study, the per-site audit trail is the natural ledger for splitting credit among datasets, institutions, and the stewards who maintain them.

What research organisations should do now

Whatever final form an S-index takes, the direction of travel is clear: funders intend to start measuring the quality and impact of data sharing, not merely its occurrence. Data custodians can prepare without waiting for the metric to be finalised. Instrument reuse now — ensure every access and analysis against your datasets is logged with enough context to demonstrate downstream impact later. Invest in metadata and harmonisation, because annotation quality appears in every serious proposal for the metric. Assign persistent identifiers to datasets and insist on data citation in any collaboration or access agreement. And if your data is too sensitive to centralise, choose infrastructure that produces evidence of reuse as a by-product of governance — so that responsible sharing is finally visible to the systems that allocate credit.

Frequently asked questions

What is the S-index?

A proposed metric for crediting researchers and institutions for high-quality, reusable data sharing — analogous to the h-index for publications. NIH is running a $1 million open-innovation competition to design and implement one.

Is the S-index an official NIH policy?

No. It is the subject of a prize competition, not a regulation. The competition’s two-phase structure — proof-of-concept winners followed by a funded implementation phase — signals serious intent, but no S-index is currently attached to funding decisions.

How would an S-index differ from simply counting dataset downloads?

Downloads measure access, not impact, and they do not exist at all for data held in governed environments. Serious S-index designs weigh metadata quality, standards alignment, timeliness, and evidenced downstream reuse — publications, derived datasets, and analyses — rather than raw access counts.

Does the S-index disadvantage sensitive data that cannot be openly shared?

It could, if designed around open-repository signals only. The corrective is to treat governed, federated reuse as first-class evidence: audit trails from Trusted Research Environments prove reuse happened without exposing any records, so data can score well because data never leaves the source, not despite it.

Who should receive S-index credit for a shared dataset?

One of the competition’s open design questions. A defensible model credits the whole chain — the principal investigator, the data stewards and curators who made the data reusable, and the institution that governs it — rather than the PI alone.

What can a data custodian do today to prepare?

Log reuse events with context, harmonise to community data models, assign persistent identifiers, require data citation in access agreements, and keep provenance from raw data through to published outputs. All of these strengthen your position under any plausible final metric.