Skip to content

Single-cell foundation models vs established baselines ​

A Liatir Scientific Showcase · completed 5 October 2026

Pretrained single-cell foundation models did not show a uniform advantage over established baselines in this study. scGPT was competitive on PBMC, while PCA, Harmony and scVI remained particularly strong on Pancreas. The result depends on the dataset and the property being measured; there is no overall winner claim.

Study source, full results and evidence · Reproducibility archive and citation

Scientific question ​

Single-cell RNA sequencing measures gene activity in individual cells. An embedding is a compact list of numbers representing each cell, used to compare cells, discover groups and predict cell types. A foundation model learns such representations from a large collection of data before this evaluation.

Do pretrained embeddings preserve biological cell types and reduce differences between experimental batches better than established methods, and what resources do they require? A batch is a group of cells measured in the same experiment or with the same technology. Mixing batches is useful only if biological distinctions are preserved.

Datasets and methods ​

DatasetCellsGenesCell typesBatchesSource
PBMC, peripheral blood mononuclear cells11,9903,34692scvi-tools 1.3.3 PBMC loader; original 10x PBMC8k + PBMC4k and scVI annotations
Pancreas16,38219,093149 technologiesscIB integration benchmark; official scvi-tools tutorial source, Figshare file 24539828

Source counts, cell/gene order and a seed-23 stratified 80/20 label split were fixed across methods. A seed fixes the random choices; stratification keeps cell types represented in both partitions. All cells remained in the common evaluation. Exact downloads, hashes, gene mappings and exclusions are recorded in the dataset manifest.

MethodRepresentation used
PCAPrincipal component analysis: a compact linear representation of 2,000 highly variable genes, using 50 components.
HarmonyBatch correction of the same PCA representation, using the recorded batch labels.
scVIA probabilistic model trained on these datasets, with 30 latent dimensions and 100 epochs.
Geneformer V1 10MPretrained cell embeddings, without fine-tuning on these datasets.
scGPT whole-humanPretrained cell embeddings, without fine-tuning on these datasets.
UCE 4-layerRequested pretrained method; both configurations blocked and unscored.

The five feasible methods completed on both datasets: ten measured configurations. scVI learned from all evaluation counts; PCA and Harmony also used all cells without test-label supervision. Only the classification training partition supplied labels to the classifiers. This tests hidden-label prediction within the datasets, rather than transfer to an unseen study or batch.

The frozen protocol specifies preprocessing, model inputs, clustering, classifiers and resource limits. Scientific parameters were fixed before scores were inspected.

Biological results ​

Logistic macro-F1 measures cell-type prediction with a fitted classifier, giving each cell type equal weight; 1 is perfect. ARI, adjusted Rand index, measures agreement between discovered groups and known cell types, with chance agreement accounted for. They measure different properties.

MethodPBMC logistic macro-F1PBMC ARIPancreas logistic macro-F1Pancreas ARI
PCA0.928800.753180.977880.71291
Harmony0.932410.706190.949530.92191
scVI0.949160.541800.971840.52441
Geneformer0.839460.289060.764350.05019
scGPT0.956850.669260.899870.27478

On PBMC, scGPT had logistic macro-F1 0.95685, compared with PCA 0.92880 and scVI 0.94916. PCA had higher clustering agreement than scGPT. On Pancreas, PCA and scVI had logistic macro-F1 0.97788 and 0.97184; scGPT had 0.89987 and Geneformer 0.76435. Harmony had ARI 0.92191. These are observations from the fixed split, without uncertainty estimates.

The full report also includes nearest-neighbor classification, mutual information between groups and labels, and cell-type separation. All values at full precision, including UCE's null scores and causes, are in the twelve-row results table.

Batch results ​

Scaled iLISI measures local mixing of batches on a zero-to-one scale, with higher values indicating more mixing. Graph connectivity measures how much of each cell type remains connected in the graph of neighboring cells. Neither alone establishes a useful biological representation.

MethodPBMC scaled iLISIPBMC connectivityPancreas scaled iLISIPancreas connectivity
PCA0.786810.773560.003090.94257
Harmony0.826470.779010.210060.88707
scVI0.761600.937850.118180.93658
Geneformer0.700820.764770.014330.50387
scGPT0.744290.847800.006370.76947

Harmony showed stronger local batch mixing than the completed pretrained methods on both datasets by this measure. scVI's PBMC connectivity was 0.93785. Batch composition and geometry affect these scores; the full measured report also reports batch separation and mixing within cell types. No combined ranking was constructed.

Inspecting the representations ​

UMAP compresses the cell representations into two dimensions for inspection. Each dot is a cell and its color is the supplied cell type. Scores were calculated on the original representations, not on these pictures. Coordinates and apparent distances between separate UMAP plots are not directly comparable. The figures below are unchanged copies of the validated study figures.

PBMC scGPT UMAP, with cells colored by the nine supplied cell types.

PBMC / scGPT: a visual view of the representation whose logistic macro-F1 was 0.95685. The plot does not establish superiority across the other metrics.

Pancreas Harmony UMAP, with cells colored by the fourteen supplied cell types.

Pancreas / Harmony: the corresponding full-representation ARI was 0.92191.

Pancreas scGPT UMAP, with cells colored by the same fourteen cell types.

Pancreas / scGPT: the corresponding ARI was 0.27478. The measured biological and batch scores support interpretation alongside this visual inspection.

All 22 original figures include both cell-type and batch views for every completed configuration.

Computational measurements ​

Nine representations were produced on an Apple M1 with 16 GiB RAM, using CPU, one numerical thread, one-cell model batches and a 2 GiB process-family memory ceiling. Those saved representations and their measurements were reused byte for byte. Only the unfinished Pancreas scGPT representation was completed on an Intel i7-8700K / NVIDIA RTX 4060 Ti, through Linux Liatir in WSL2, the Windows subsystem for running Linux. That run used six threads, batches of 16, a 16 GiB workload RAM ceiling and a 6 GiB whole-GPU memory ceiling.

Representation time includes imports, preprocessing, model loading, inference or training, and serialization. Downloads, installation, common evaluation and UMAP are excluded. RAM below is the worker process's operating system high-water mark; GiB means 1,024³ bytes.

DatasetMethodProducerSecondsPeak worker RAM, GiB
PBMCPCAMac / CPU7.511.192
PBMCHarmonyMac / CPU15.201.210
PBMCscVIMac / CPU249.220.949
PBMCGeneformerMac / CPU804.100.681
PBMCscGPTMac / CPU4,994.071.153
PancreasPCAMac / CPU7.941.799
PancreasHarmonyMac / CPU22.931.714
PancreasscVIMac / CPU333.291.674
PancreasGeneformerMac / CPU8,448.011.715
PancreasscGPTPC / WSL2 / CUDA515.542.346

Pancreas scGPT's GPU tensor peak was 3.269 GiB; the independent whole-device monitor measured 4,799,332,352 bytes, including graphics and other processes. The workload had zero swapping. CPU accelerator measurements are null, with their reasons recorded. Historical disk figures count file lengths through symlink aliases and are not physical disk usage; the separate PC disk audit retains that distinction.

Logistic macro-F1 versus representation runtime, with separate Pancreas and PBMC panels and hardware labels.

The runtime axis is logarithmic. Pancreas scGPT uses different hardware and resource limits from the other points, so this is not a same-host speed ranking. The full compute table also preserves model identities, embedding sizes and measurement scopes.

Interpretation and limitations ​

Pretrained representations were useful in some measured settings, especially scGPT cell-type prediction on PBMC. The observations do not support a uniform advantage over established approaches. Biological preservation, batch mixing and resource use should be assessed separately for the intended analysis.

  • Two datasets, one seed. There are no confidence intervals or statistical superiority claims. Rare Pancreas labels further limit interpretation.
  • Within-dataset evaluation. scVI was trained on the evaluation datasets; unsupervised fitting used all cells. The label split is not a held-out study or batch. Possible overlap with foundation-model pretraining data was not ruled out.
  • Input constraints. PBMC uses a historically selected 3,346-gene subset. Pancreas contains fractional quantification values, preserved without rounding; integer count-model assumptions are imperfect. Each pretrained model uses its required vocabulary and input transformations.
  • Heterogeneous hardware. The Mac and PC costs cannot establish a same-host performance ranking. RAM and GPU measurements have different scopes. Historical interrupted work is retained separately from the fresh complete GPU runtime.
  • UCE has no scientific scores. On the original Mac, its checkpoint tensors exceeded the approved 2 GiB budget. On the PC, no signed Linux UCE target was published. These are execution blockers, not evidence of poor representations or inability to run with more resources.

The complete limitations and failed/interrupted attempts remain in the technical package.

What Liatir handled ​

The study ran through Tools → Single-cell study in the compiled desktop app. Liatir handled model installation through AI Models, dataset preparation, the study's analysis steps, resource safeguards, logs and progress in Jobs, and measurements, charts and export in Results. Saved runs had their own workspace and run identities. Reuse required matching data, split, code and embedding checksums, which identify the exact saved contents.

Before the complete GPU run, real interruption/resume produced exactly equal vectors, and incompatible identities and impossible resource limits were refused. The final verification used a native Linux app under WSL2 with a virtual display and a development frontend; it does not certify a Windows release package.

Independent evaluation reproduction matched all completed biological and batch scores at absolute tolerance 1e-9 and relative tolerance 1e-8. Original source-count checks passed exactly, including all fractional Pancreas values. The historical repository, native UI and scientific regression checks are detailed in tested reproduction and execution evidence.

Reproduce and cite ​

Use Liatir's visual interface to repeat the study:

  1. Choose a workspace and make the required models available through AI Models.
  2. Open the single-cell study in Tools. Select PBMC and the study's six methods: PCA, Harmony, scVI, Geneformer, scGPT and UCE.
  3. Use the complete dataset rather than the small-sample stability check, and choose an execution setting supported by your computer.
  4. Start the study and follow progress and diagnostics in Jobs.
  5. Inspect the measurements, figures and any blocked methods in Results, then export the study.
  6. Repeat for Pancreas with the same scientific settings.

Liatir handles the preparation and execution; initial downloads require an internet connection. Unavailable models and resource limits remain explicit blockers. New runs record their own model versions, hardware and execution costs. Continuing a saved study reuses verified completed work and preserves its original history.

Reproducing the study in Liatir

GitHub tracks the protocol, source, dataset/provenance manifests, small result tables, all figures and validation evidence. Zenodo holds the complete large reproducibility bundle, approximately 1.8 GB, including data and saved results that are deliberately outside Git. Download the original study artifacts and cite them through 10.5281/zenodo.23187931. Use the record's citation metadata when citing the archive.