bioRxiv · 10.64898/2026.06.18.733285
Systematic benchmarking of zero-shot utility and robustness in single-cell transcriptomic foundation models
Abstract
Single-cell foundation models (scFMs) have emerged as powerful representation learning approaches for single-cell transcriptomics. However, the utility and robustness of their pretrained representations across diverse analytical tasks and data conditions remain insufficiently characterized, particularly in zero-shot settings without task-specific fine-tuning. Here, we systematically analyze zero-shot performance of single-cell transcriptomic representations across 20 methods, 6 downstream tasks and 1,607 datasets comprising nearly 21.8 million cells. We evaluate model behavior along three complementary dimensions: utility on original datasets, robustness to controlled changes in dataset structure, and exploratory associations between dataset characteristics and performance variation. Our results show that scFM performance is strongly task dependent, with no single method consistently outperforming others across cell- and gene-level analyses. Notably, high utility on original datasets did not necessarily translate into robustness under structural perturbations, and several top-ranking methods were sensitive to changes in cell number, gene number, class composition, class imbalance, and batch complexity. Conventional statistical and task-specific methods remained competitive in several settings, while greater computational cost did not consistently correspond to better performance. Driver analyses further identified task-specific associations between performance and dataset characteristics, including cell-type complexity, train-test class overlap, batch number, and regulatory target-set size. Together, these findings show that the zero-shot utility and robustness of pretrained scFM representations depend jointly on analytical task and dataset structure. Our study provides a practical basis for context-aware representation selection and underscores the importance of evaluating structural robustness alongside utility when developing and applying scFMs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Liu, T., Feng, T., Pan, X., Chen, Y., Ren, L., Ye, X., Sakurai, T., Lin, H., Zhang, Y.. 2026-06-23. Systematic benchmarking of zero-shot utility and robustness in single-cell transcriptomic foundation models. https://doi.org/10.64898/2026.06.18.733285
Cite the original work for its findings. Save a collection to share your selection of sources.