bioRxiv · 10.64898/2026.04.01.714123
CellBench-LS: Benchmark Evaluation of Single-cell Foundation Models for Low-supervision Scenarios
Abstract
Single-cell foundation models are trained on millions of cells, but their value when downstream labels are scarce remains uncertain. Here we compare seven such models with three established representations across clustering, batch correction, cell type annotation, gene expression reconstruction and perturbation prediction. Frozen representations are tested directly for clustering and batch correction and with lightweight task heads for few-shot prediction. CellPLM led the aggregate clustering, annotation and batch-correction comparisons, scMulan led perturbation prediction, and principal component analysis led reconstruction. Of the seven foundation models, one exceeded the strongest classical baseline in clustering, three in batch correction, four in annotation, none in reconstruction and six in perturbation prediction. Representation analyses linked reconstruction performance to preservation of linear gene relationships, but cross-model gene-relation agreement did not exceed a dimension-matched random-projection null. These results show that the utility of a cell representation depends on the information required by the downstream task.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xu, Y., Li, Y., Yuan, Y., Yu, C., Zang, Z.. 2026-04-05. CellBench-LS: Benchmark Evaluation of Single-cell Foundation Models for Low-supervision Scenarios. https://doi.org/10.64898/2026.04.01.714123
Cite the original work for its findings. Save a collection to share your selection of sources.