Lifebit logo
BlogUncategorizedPolygenic Risk Scores on Federated Data: Build and Validate

Polygenic Risk Scores on Federated Data: Build and Validate

Intricate dark blue abstract grid texture with flowing wavy pattern, perfect for backgrounds.
Photo by 3D Render on Pexels

A polygenic risk score (PRS) sums the small effects of many genetic variants into a single estimate of a person’s inherited predisposition to a disease or trait. On federated data, researchers build and validate a PRS by fixing the score’s variant weights from published genome-wide association summary statistics, running the same scoring pipeline inside each participating biobank, and bringing back only aggregate performance metrics such as area under the curve, variance explained, and calibration. Because individual genotypes stay where they are held, federation lets a PRS be tested across many cohorts and ancestries that could never legally be pooled.

Why federated PRS matters now

Polygenic scores have moved from research papers toward clinical pilots in cardiovascular disease, breast cancer, and diabetes screening. The PGS Catalog, an open database of published scores, now holds thousands of scores for hundreds of traits. The bottleneck is no longer building scores. It is validating them in populations that look like the patients who will actually receive them.

That bottleneck is well documented. Most genome-wide association studies (GWAS) have been conducted in people of European ancestry, and a widely cited 2019 analysis by Martin and colleagues in Nature Genetics showed that PRS accuracy drops substantially when scores are applied to other ancestries. A score validated in one population can mislead clinicians in another. Fixing that requires testing scores across diverse cohorts held by different institutions in different countries.

Those cohorts are exactly the ones that cannot be pooled. Genomic data is special category data under the General Data Protection Regulation (GDPR), many countries impose data residency rules on genomes, and consent terms often restrict transfer. Federated analysis is the practical way to test one score against many populations.

How PRS construction works

Start from summary statistics

Most scores are built from GWAS summary statistics: for each variant, an effect size, standard error, and p-value. These aggregate statistics are routinely published and do not contain individual genotypes, which makes them a natural starting point for federated work.

Choose a weighting method

The simplest approach is clumping and thresholding (C+T), which keeps the strongest independent variants below a p-value cut-off. Bayesian methods such as LDpred2 and PRS-CS adjust effect sizes for linkage disequilibrium, the correlation between nearby variants, and usually improve accuracy. Penalized regression methods such as lassosum take a similar approach. Each method needs a linkage disequilibrium reference panel that matches the ancestry of the target cohort as closely as possible.

Tune, then freeze

Methods with tuning parameters, such as the p-value threshold or a shrinkage prior, need a tuning cohort separate from the final test cohort. Once parameters are chosen, the weights are frozen. A frozen weight file is simply a list of variants, effect alleles, and weights, and it can be distributed to every site without privacy concerns.

The federation angle: fixed weights travel, genotypes stay

PRS validation is an unusually good fit for federation because the expensive, sensitive step (computing each person’s score) is local and deterministic, while the useful output (how well the score performs) is aggregate. In a federated Trusted Research Environment (TRE), the frozen weight file and a containerized scoring pipeline are sent to each biobank. The pipeline computes scores locally, fits the evaluation models against local phenotypes, and returns metrics. Throughout, data never leaves the source, and each result passes an airlock review before release.

Federation also improves scientific quality. Each site keeps control of its own quality control, ancestry assignment, and phenotype definitions, but runs them through a common, versioned pipeline, so results are comparable. When phenotype definitions are harmonized to a common data model such as OMOP, the same case definition for coronary artery disease or type 2 diabetes runs identically across sites. The result is a set of per-cohort, per-ancestry performance estimates that can be meta-analyzed, which is exactly what reviewers and clinical guideline groups ask for.

A step-by-step guide to building and validating a PRS on federated data

  1. Define the target and population. State the disease or trait, the intended use (research stratification, screening, or clinical decision support), and the ancestries in which the score must perform.
  2. Select discovery summary statistics. Choose GWAS results with no sample overlap with any validation cohort. Overlap inflates accuracy and is one of the most common errors in PRS papers.
  3. Harmonize variants. Align genome builds, strand, and effect alleles, and resolve ambiguous palindromic variants. Publish the harmonized weight file with a version identifier.
  4. Tune in one federated site, test in others. Use a designated tuning cohort to set parameters, then freeze the weights before any test cohort is scored.
  5. Ship a containerized pipeline. Package genotype quality control, imputation checks, scoring, and evaluation into one versioned container so every site runs identical code.
  6. Return aggregate metrics only. Each site reports discrimination, variance explained, calibration, and effect sizes by ancestry group, with small cells suppressed by disclosure control.
  7. Meta-analyze and report. Combine site results with random-effects meta-analysis and report using the PRS Reporting Standards (PRS-RS) published by the ClinGen Complex Disease Working Group and the Polygenic Risk Score Task Force of the International Common Disease Alliance.

Handling ancestry in a federated network

Ancestry assignment should happen at each site using a shared method, for example projecting samples onto principal components computed from a public reference panel such as the 1000 Genomes Project. Using the same reference and the same projection code everywhere means that “African ancestry” or “South Asian ancestry” groups are defined consistently, which is essential when results by group are later combined. Where a site has too few participants in a group to report safely, that group should be suppressed at that site and captured through the meta-analysis instead.

Which validation metrics to return

MetricWhat it measuresWhen to use itDisclosure consideration
Area under the ROC curve (AUC)Discrimination for binary outcomesDisease risk scoresLow risk as a single aggregate; report confidence intervals
Incremental AUCAdded value over age, sex, and clinical factorsAny score proposed for clinical useLow risk
Variance explained (R squared, or liability-scale R squared)Proportion of trait variation capturedQuantitative traits and heritability comparisonsLow risk
Odds or hazard ratio per standard deviationEffect size of the scoreComparing scores across cohortsLow risk
Top versus rest comparisonRisk in the highest score percentilesScreening and stratification use casesCheck cell counts in small cohorts
Calibration (observed versus predicted by decile)Whether absolute risk estimates are correctAbsolute risk modelsDecile tables can contain small counts; apply suppression

Discrimination alone is not enough. A score can rank people well and still overstate or understate absolute risk in a new population. Calibration results by ancestry are what tell a health system whether the score can be used as is or needs recalibration locally.

Real-world context

Large public biobanks have defined how PRS research is done. UK Biobank has been the discovery or validation cohort for a large share of published scores, and the All of Us Research Program in the United States was designed to recruit a diverse participant base, which makes it important for cross-ancestry evaluation. The PGS Catalog, maintained as an open resource, records the development and evaluation cohorts of each score so users can see where it has and has not been tested.

National genomics programs are where federated validation becomes routine. Genomics England, a Lifebit customer, holds whole-genome data linked to clinical records inside a controlled research environment. Programs of this type can evaluate a published score against their own participants without exporting genotypes, and contribute aggregate results to wider meta-analyses. The same pattern extends across borders: a score tuned in one country’s biobank can be tested in another’s without either dataset moving.

Common pitfalls

Sample overlap

If validation participants were part of the discovery GWAS, performance is inflated. Federated programs should check overlap at the cohort level before scoring, since individual-level checks across sites are not possible without linkage.

Population stratification

Score distributions shift with ancestry. Always adjust evaluation models for principal components of ancestry, and report results within ancestry groups rather than only pooled.

Inconsistent phenotypes

If one site defines a case by a single diagnosis code and another requires two, differences in performance may reflect definitions, not genetics. Harmonized, versioned phenotype definitions remove this source of noise.

Retuning in the test set

Adjusting thresholds after seeing test results is overfitting. Freeze the weights and the evaluation plan before any test site runs the pipeline.

Ignoring genotyping platform differences

Sites that genotype on different arrays or use different imputation panels will have different variant coverage. A score whose weights include variants missing at one site will perform worse there for technical, not biological, reasons. Report the proportion of score variants available at each site alongside performance.

Releasing individual scores

Per-person scores are derived genetic data and should stay inside each environment. Only aggregate metrics should pass the airlock.

What to do next

Start by listing the cohorts your score must be validated in and checking which of them could never be pooled. That list usually makes the case for federation on its own. Lifebit’s guide to federated genomic analysis at population scale covers the operational side of running pipelines across biobanks, and the article on output checking and statistical disclosure control explains how aggregate metrics are reviewed before release. For the underlying platform, see the overview of the federated Trusted Research Environment.

Frequently asked questions

What is a polygenic risk score?

A polygenic risk score is a weighted sum of many genetic variants, each with a small effect, that estimates a person’s inherited predisposition to a disease or trait relative to others in the same population.

Can a polygenic risk score be validated without pooling genetic data?

Yes. The frozen variant weights and a scoring pipeline are sent to each biobank, scores are computed locally, and only aggregate metrics such as AUC, variance explained, and calibration are returned and meta-analyzed.

Why do polygenic risk scores perform worse in some ancestries?

Most GWAS have been run in people of European ancestry, and differences in allele frequencies and linkage disequilibrium between populations reduce how well those effect estimates transfer. Validation in diverse cohorts is needed to measure and correct this.

Which PRS methods work best in federated settings?

Methods that produce a fixed weight file, such as clumping and thresholding, LDpred2, and PRS-CS, work well because the weights can be distributed freely and applied identically at each site. The key is to tune in one cohort and freeze weights before testing.

What metrics should be reported when validating a PRS?

Report discrimination (AUC or incremental AUC), variance explained, effect size per standard deviation, and calibration, stratified by ancestry. The PRS Reporting Standards provide a checklist for publication.

Are individual polygenic scores considered sensitive data?

Yes. Individual scores are derived from genetic data and can reveal health risks, so they should remain inside the research environment. Only aggregate results should be released after disclosure review.