Lifebit logo
BlogUncategorizedPopulation-Scale Proteomics: Data Infrastructure Guide

Population-Scale Proteomics: Data Infrastructure Guide

Detailed close-up of ethernet cables and network connections on a router, showcasing modern technology.
Photo by Pixabay on Pexels

Population-scale proteomics needs data infrastructure that can store and version protein measurements for hundreds of thousands of people, harmonize results across assay platforms, link them to genomic and clinical records, and let approved researchers analyze everything without copying it out. In practice that means a federated Trusted Research Environment (TRE) with platform-aware harmonization, scalable compute for protein quantitative trait loci and association studies, and automated output checking.

Why proteomics infrastructure matters now

Proteomics is the large-scale measurement of proteins, the molecules that carry out most of the work in cells and that most drugs target. For a long time, population cohorts measured a handful of proteins per participant. High-throughput affinity platforms changed that. Proximity extension assays, used by Olink, and aptamer-based assays, used by SomaLogic’s SomaScan, now measure thousands of plasma proteins from a small blood sample.

The UK Biobank Pharma Proteomics Project is the clearest public example of the shift. Its first phase, published in Nature in 2023, reported plasma levels of nearly 3,000 proteins in about 54,000 UK Biobank participants, and linked them to genetic variants and disease outcomes. A larger follow-on phase has since been announced to profile many more samples. Other population studies, including cohorts in Iceland, the UK, and elsewhere, have produced large aptamer and mass spectrometry datasets of their own.

These datasets are valuable to pharma because a protein that is genetically linked to disease is a strong candidate drug target. That value brings pressure to share widely, and the May 2026 UK Biobank incident showed the weakness of access models that rely on researchers downloading data and behaving well afterward. Proteomics adds new identifiable signals to already sensitive cohorts, so the infrastructure question is not optional.

What makes proteomics data different

Platform-specific measurements

Proteomics values are not directly comparable across technologies. Olink reports Normalized Protein eXpression (NPX), a relative value on a log2 scale. SomaScan reports relative fluorescence units that are then normalized. Mass spectrometry produces spectra that are processed into peptide and protein abundances, often stored in open formats such as mzML defined by the Human Proteome Organization Proteomics Standards Initiative. Two platforms can target the same protein and still disagree, because the assays bind different parts of the protein or detect different forms of it.

Batch and plate effects

Samples are run on plates, in batches, over months or years. Plate effects, reagent lot changes, and sample handling (such as time to freeze) all shift values. Any serious infrastructure has to store plate and batch identifiers, bridging sample results, and quality control flags alongside the protein values, or later analyses will confuse lab artifacts with biology.

Deep linkage to genomics and health records

The scientific payoff comes from linkage. Protein quantitative trait loci (pQTL) studies connect genetic variants to protein levels. Association studies connect protein levels to diagnoses, prescriptions, and outcomes recorded in electronic health records. That makes proteomics one layer of a multi-omic cohort, not a standalone file, and it inherits the re-identification risk of every layer it is linked to.

The federated TRE angle

A federated TRE keeps each dataset where its custodian holds it and brings approved analysis to the data. Data never leaves the source. For proteomics this matters for three reasons.

First, linkage stays inside the boundary. Proteomics, genotypes, and clinical data can be joined for analysis within the TRE without assembling a combined, exportable file that would be very hard to protect once it existed.

Second, multi-cohort studies become practical. A pharma sponsor or consortium can run the same pQTL or association workflow across several biobanks, each in its own environment, and combine summary statistics through meta-analysis. This avoids negotiating a separate data transfer for every cohort. The mechanics of this approach are covered in our guide to federated genomic analysis at population scale, and the same pattern applies to proteins.

Third, outputs are checked before release. Lifebit’s federated TRE routes exported results through an automated airlock, which applies disclosure rules such as minimum cell counts and flags unusual outputs for human review. Summary statistics from a well-powered study are usually low risk. Per-participant protein values are not, and should not leave the environment.

Governance for commercial and academic access

Proteomics projects are often funded by industry consortia, with an agreed period of exclusive access before data opens to the wider research community. Infrastructure has to enforce those terms precisely. In a federated TRE, access is granted per project and per data release, so a consortium member can work on its licensed release while academic researchers see only what their approval allows. When the exclusive period ends, the custodian changes permissions rather than distributing new copies. The same controls support participant consent: if a participant withdraws, their data is removed at the source, and no stale copies remain in researcher environments.

A practical infrastructure framework

The table below compares how a typical download-based model and a federated TRE handle the main requirements of population-scale proteomics.

RequirementDownload or centralized SaaS modelFederated TRE (Lifebit)
Storage of protein valuesCopied into researcher or vendor environmentsStored once at the custodian, versioned by data release
Linkage to genotypes and health recordsLinked files often exported togetherLinked inside the TRE; combined files never exported
Cross-platform harmonizationEach research team repeats the workDone once by the custodian, published as curated variables
Multi-cohort studiesSeparate transfer agreement per cohortSame workflow runs at each site; summary results combined
Output controlRelies on researcher compliance after downloadAirlock checks every exported result
AuditLimited visibility once data leavesFull record of who ran what, on which data release

1. Model the data properly

Store protein values with their metadata: assay platform and version, panel, plate, batch, sample collection and processing times, and quality flags such as values below the limit of detection. Keep raw and normalized values separately so researchers can reprocess if normalization methods change.

2. Harmonize across platforms and releases

Map every assay target to a stable protein identifier, such as a UniProt accession, and record the assay’s own identifier alongside it. Publish documented cross-platform comparisons so researchers know which proteins agree well between Olink and SomaScan and which do not. Link clinical variables to a common model such as the OMOP Common Data Model, so proteomic associations use consistent outcome definitions. Our multi-omics harmonization guide goes into this step in more depth.

3. Provide compute sized for the workload

A genome-wide pQTL scan across thousands of proteins is thousands of genome-wide association studies. The TRE needs elastic batch compute and containerized workflows, not just an interactive notebook, or researchers will try to move data somewhere faster.

4. Set output rules for proteomics

Agree in advance what can leave: summary statistics above a minimum sample size, aggregate plots, and model coefficients. Individual-level values, small subgroups, and outputs that could be combined to reconstruct individual data stay inside.

5. Plan for data releases

Proteomics datasets grow in phases. Version every release, keep earlier versions reproducible, and record which release each analysis used, so published results can be traced and repeated.

6. Record provenance for derived variables

Researchers often create derived variables, such as protein risk scores or residualized values adjusted for age, sex, and batch. Keep a record of the code, inputs, and data release behind each one. When a derived variable proves useful, the custodian can review it and return it to the curated dataset, so the next team starts from a documented, shared version rather than rebuilding it.

Real-world example

UK Biobank’s proteomics data is available to approved researchers through its cloud-based Research Analysis Platform rather than as a bulk download, which reflects the broader move toward analysis in a controlled environment. Federation extends that principle across institutions. National and regional programs that already run a federated TRE for genomic and clinical data, such as those supported by Lifebit for Genomics England and for Canada’s CanPath cohort, can add proteomics as another linked layer without creating a new sharing route. The design principles for that kind of national setup are described in our federated TRE reference architecture.

Common pitfalls and objections

Treating protein data as low risk

Protein levels can reveal sex, pregnancy, disease status, and medication use, and they are linked to genotypes that are inherently identifying. Treat per-participant proteomics as sensitive health data.

Pooling platforms without harmonization

Combining Olink and SomaScan values as if they were the same measurement produces misleading results. Keep platform as an explicit variable and publish guidance on comparability.

Ignoring pre-analytical metadata

If sample handling times and batches are missing, analysts cannot separate biology from lab effects. Capture this metadata at ingestion; it is very hard to recover later.

Assuming federation is too slow for pQTL work

Performance depends on compute provisioning. Running a large association workflow next to the data avoids moving large files, and summary statistics are small enough to combine quickly.

What to do next

If you are planning or expanding a proteomics study, start with the data model and the output rules, then size the compute. Decide early how proteomics will be linked to genotypes and clinical records, and make sure that linkage happens inside a controlled environment. The foundation for that is a federated Trusted Research Environment designed to hold multiple omics layers under one governance model, where custodians keep control and researchers still get the analytical depth they need.

Frequently asked questions

What is population-scale proteomics?

It is the measurement of hundreds or thousands of proteins, usually in blood plasma, across very large cohorts of people. The goal is to find proteins linked to genetics, disease risk, and drug response at a scale that gives robust statistical power.

What is the difference between Olink and SomaScan data?

Olink uses proximity extension assays and reports Normalized Protein eXpression (NPX) on a log2 scale. SomaScan uses aptamers and reports normalized relative fluorescence units. They can target the same protein but measure it differently, so values are not directly interchangeable.

What is a pQTL?

A protein quantitative trait locus (pQTL) is a genetic variant associated with the level of a protein. pQTL studies help identify causal links between genes, proteins, and disease, which is why they are valuable for drug target discovery.

Is proteomics data identifiable?

Per-participant proteomics can reveal sex, disease, and medication use, and it is usually linked to genotypes and health records. It should be handled as sensitive health data, with only checked aggregate results released.

Why use a federated TRE for proteomics?

A federated TRE keeps proteomics, genomics, and clinical data at the custodian, lets approved researchers analyze the linked data in place, and checks every exported result. It also allows the same analysis to run across several cohorts without transferring data between them.

How should proteomics data be versioned?

Each data release should have a fixed version with its platform, normalization method, and quality control rules recorded. Analyses should log which release they used, so results can be reproduced when new phases of data arrive.