Lifebit logo
BlogUncategorizedSingle-Cell and Spatial Omics Data in a TRE

Single-Cell and Spatial Omics Data in a TRE

Modern abstract geometric art with dynamic blue and gray structures creating a striking visual impact.
Photo by Steve A Johnson on Pexels

Single-cell and spatial omics data belong in a Trusted Research Environment (TRE) because the files are both identifying and enormous: sequencing reads carry the donor’s genotype, and tissue images can reveal clinical context. A federated TRE lets approved researchers run standard single-cell and spatial pipelines where the data is stored, then release only checked, aggregate results, so the raw reads and images never have to be copied to a researcher’s laptop or a shared cloud bucket.

Why single-cell and spatial data now need a secure home

Single-cell RNA sequencing (scRNA-seq) measures gene expression in individual cells rather than averaging across a tissue sample. Spatial transcriptomics goes a step further and records where in a tissue section each measurement was taken. Together they have moved from specialist labs into mainstream disease research, drug target discovery, and large atlas efforts such as the Human Cell Atlas, a public international consortium mapping every cell type in the human body.

That growth has exposed a governance gap. For years, many teams treated processed single-cell matrices as low risk, because a table of cells by genes does not look like a patient record. The underlying data tells a different story. Raw reads from scRNA-seq overlap common genetic variants, and the same property that lets donor demultiplexing tools assign pooled cells back to individual donors also means a determined analyst could link a dataset to a person. Spatial assays add high-resolution tissue images, often from rare tumors or pediatric samples, where the diagnosis itself narrows the pool of possible donors.

The May 2026 UK Biobank incident showed what happens when governance depends on researchers behaving well after they have downloaded data. Approved users moved data out through a normal workflow, with no policy broken. For single-cell and spatial collections, where a single study can combine genotype-bearing reads, histology, and clinical labels, the same failure would be hard to reverse.

What makes these data types hard to govern

Volume and file structure

A single-cell study produces several layers of data. Raw FASTQ files from the sequencer are the largest and most identifying. Alignment and counting tools such as Cell Ranger produce BAM files and sparse count matrices. Downstream, analysts work with AnnData (.h5ad) files in Python’s Scanpy ecosystem or Seurat objects in R. Spatial platforms add image files, often OME-TIFF or OME-Zarr, plus coordinate tables that tie each spot or cell back to a position in the tissue.

Each layer has a different risk profile and a different audience. Methods developers often want raw reads. Biologists usually need only the annotated matrix. A clinical collaborator may only need cell-type proportions per sample. A good TRE design gives each group the layer it needs, inside the environment, instead of one download that contains everything.

Re-identification risk hidden in “processed” data

Processed matrices are not automatically safe. Expression of sex-linked genes reveals donor sex. Rare cell states can point to a rare condition. When a study includes a small number of donors, sample-level metadata such as age, site, and diagnosis may be enough to single someone out. This is the same problem that statistical disclosure control addresses for tabular health data, and the same principles apply. Our guide to output checking and statistical disclosure control in TREs covers the checks in detail.

Metadata inconsistency across labs

Different labs annotate cell types with different names, use different tissue dissociation protocols, and record donor metadata in free text. Combining datasets across institutions requires harmonized ontologies, such as the Cell Ontology for cell types and Uberon for anatomy, before any cross-study comparison means anything.

The federated TRE approach

A federated TRE reverses the usual flow. Instead of pooling single-cell data in one central repository, each data custodian keeps its data in its own environment, and analysis is sent to where the data lives. Data never leaves the source. Researchers see results, not raw files.

For single-cell and spatial work, this has three practical effects. First, the largest files stay put, which avoids the cost and delay of moving raw sequencing data between clouds or countries. Second, genotype-bearing reads remain under the custodian’s control, so a hospital or biobank can share a spatial tumor atlas without handing over the reads that could identify patients. Third, every result that leaves the environment passes through an airlock, where figures, tables, and model outputs are checked for disclosure risk before release. Lifebit’s federated TRE applies automated airlock checks to outputs, with human review for anything the rules flag. For more on how these checks work, see what an automated airlock is in a TRE.

This model also suits multi-site atlas projects. A consortium can agree a common processing pipeline and a harmonized annotation scheme, run the same containerized workflow at every site, and combine summary outputs or model parameters rather than raw cells.

A practical framework for running single-cell and spatial omics in a TRE

The table below maps each data layer to where it should live and who should see it. It is a starting point for a data access policy, not a fixed rule, and custodians should adjust it to their own consent terms and legal basis.

Data layerTypical formatIdentifiabilityRecommended handling in a federated TRE
Raw sequencing readsFASTQ, BAM/CRAMHigh (contains genotype)Stays at source; accessible only to approved methods workflows; never exported
Count matricesMTX, HDF5, .h5ad, Seurat objectModerate (sex, rare states, small donor groups)Analyzed inside the TRE; export limited to aggregates that pass the airlock
Spatial imagesOME-TIFF, OME-ZarrModerate to high for rare tissuesViewed and analyzed in the TRE; exported figures reviewed for identifying detail
Donor and sample metadataTabular, linked to clinical recordsHigh when combined with rare diagnosesHarmonized to a common model; linked inside the TRE only
Summary resultsCell-type proportions, marker tables, figuresLow once checkedReleased through airlock with small-count suppression

Step 1: Classify each layer before ingestion

Decide which layers you hold and which your consent allows researchers to use. Many biobanks can support matrix-level analysis broadly while restricting raw reads to a narrow set of approved projects.

Step 2: Standardize the processing pipeline

Use versioned, containerized workflows so every site produces comparable outputs. Community pipelines such as the nf-core scrnaseq pipeline give a documented, reproducible baseline that can run inside a TRE without internet access to raw data.

Step 3: Harmonize annotations and metadata

Map cell types to the Cell Ontology, tissues to Uberon, and donor clinical data to a common model such as the OMOP Common Data Model where clinical linkage is needed. Without this step, cross-site comparisons compare labels rather than biology.

Step 4: Provide the right compute

Single-cell analysis is memory heavy, and spatial image analysis often needs GPUs. The TRE should offer scalable workspaces with Jupyter, RStudio, and the common libraries preinstalled, so researchers do not feel pressure to work around the environment.

Step 5: Define output rules for figures

UMAP plots, spatial heatmaps, and violin plots can all leak information when groups are small. Set minimum donor and cell counts for released figures, and check that exported images do not include slide labels or embedded metadata.

Real-world context

Public infrastructure already reflects this layered model. The Human Cell Atlas and portals such as CZ CELLxGENE publish processed matrices openly for many datasets, while raw sequencing data from human donors is typically held under controlled access in archives such as the European Genome-phenome Archive. That split recognizes that the raw layer carries the highest risk.

National genomics programs have taken the same logic further for whole-genome data. Genomics England, a Lifebit customer, gives approved researchers access to genomic and clinical data inside a secure research environment rather than distributing files. As hospitals and biobanks add single-cell and spatial assays to their cohorts, extending that same federated TRE model to new data types is simpler than building a parallel sharing process. The broader case for multi-omic federation is laid out in why federated access is the future of multi-omic data sharing.

Common pitfalls and objections

“Processed matrices are anonymous, so we can share them freely”

Sometimes they can be, especially for large donor groups with minimal metadata. The decision should follow a documented risk assessment, not an assumption. Small studies, rare diseases, and pediatric cohorts deserve particular care.

“Researchers need raw data for methods work”

Some do. A TRE can give methods developers access to raw reads inside the environment, with enough compute to run alignment and demultiplexing. What they cannot do is copy the reads out. In practice most methods work only needs the data to be reachable, not portable.

“Federation will slow analysis down”

Moving terabytes of raw sequencing data is slow too, and data transfer agreements often take longer than the analysis. Running workflows where the data sits removes the transfer step entirely. The real performance question is whether the TRE provides enough memory and GPU capacity, which is a design choice rather than a limitation of federation.

“Every site processes data differently”

That is a harmonization problem, and it exists whether data is federated or pooled. Agreeing pipelines and ontologies up front is what makes multi-site single-cell analysis reproducible in either model.

What to do next

If you hold single-cell or spatial data today, start with an inventory: which layers you store, where they sit, what your consent allows, and who currently has copies. Then compare your current sharing process against the framework above. Most custodians find that the raw layer is the main exposure and the easiest to protect by keeping it at source.

For a wider view of how a federated TRE is structured and evaluated, read our guide to Lifebit’s Federated Trusted Research Environment, then map your single-cell and spatial workflows onto it layer by layer.

Frequently asked questions

Can single-cell RNA sequencing data identify a person?

Yes, in some cases. Raw reads overlap common genetic variants, which can be matched to a donor’s genotype. Processed matrices carry lower but not zero risk, particularly for small studies, rare diseases, or datasets with detailed donor metadata.

What is spatial transcriptomics?

Spatial transcriptomics measures gene expression while keeping track of where each measurement sits in a tissue section. It combines sequencing data with tissue images, so researchers can see how cell types are arranged and interact within an organ or tumor.

Why use a Trusted Research Environment for single-cell data?

A TRE keeps sensitive files in a controlled workspace where approved researchers can analyze them, while every exported result is checked for disclosure risk. It avoids distributing copies of genotype-bearing reads and tissue images that cannot be recalled once shared.

Which file formats are common in single-cell and spatial analysis?

Raw data is usually FASTQ, aligned data BAM or CRAM, and count matrices MTX, HDF5, AnnData (.h5ad), or Seurat objects. Spatial platforms add images in formats such as OME-TIFF or OME-Zarr along with coordinate tables.

How does a federated TRE handle multi-site single-cell studies?

Each site keeps its data and runs the same containerized workflow locally. Harmonized annotations let results be compared, and only aggregate outputs or model parameters are combined, after passing airlock checks at each site.

Do researchers lose flexibility when working inside a TRE?

Not if the environment is well provisioned. Researchers can use Jupyter, RStudio, Scanpy, Seurat, and GPU compute inside the TRE. The main difference is that raw files cannot be downloaded, and exported results go through review.