bioRxiv · 10.64898/2026.06.29.735389
Raw-count embeddings improve single-cell foundation models
Abstract
Single-cell transformer foundation models have grown to hundreds of millions of parameters, yet the preprocessing choices that underlie them, including gene ranking and library-size normalisation, have not been systematically benchmarked. Testing seven strategies, we find these elaborations are largely unnecessary: non-normalised, log-transformed counts give the best performance, and gene order barely matters, with even random ordering outperforming sophisticated rank-based schemes. The resulting model, Gene Intelligence, projects log1p-transformed raw counts directly onto each token embedding and jointly predicts masked tokens and counts, using no normalisation, positional encoding, or read-depth tokens. Despite this simplicity, it achieves state-of-the-art performance in the tested gene-level tasks and in doublet detection, and matches large current foundation models on cell-classification tasks while using 10-to 200-fold fewer parameters.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Schlede, S., Muruganandan, T. P., Gojjam Kantharaju, S., Kisis, I., Boecker, M., Kim Alves Carpinteiro, M., Schmitz, A., Buchwald, L. M., Sakthivelu, V., Gülcüler Balta, G. S., Anstötz, M., Rueger, M. A., Thomas, R. K., Beleggia, F.. 2026-07-03. Raw-count embeddings improve single-cell foundation models. https://doi.org/10.64898/2026.06.29.735389
Cite the original work for its findings. Save a collection to share your selection of sources.