bioRxiv · 10.64898/2026.09.15.751768
Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics
Abstract
Models that combine tissue images with pretrained gene representations aim to predict spatial expression for genes not used to fit the downstream predictor. Yet success on held-out genes can reflect two capabilities: estimating a gene's mean expression across tissue locations and recovering its spatial variation. Across four cohorts spanning three human brain regions and HER2-positive breast cancer, we evaluated held-out genes in held-out individuals and separated these components. For spatial predictors using fixed gene representations from Decima or scGPT, reductions in gene-mean error accounted for more than 91% of the reduction in mean squared error relative to matched random vectors. Independently fitted mean-only models using the same representations but no tissue images retained 90-99% of the corresponding gain in full-matrix correlation. Spatial gains were smaller on average, increased with expression variation in training tissue and differed across cohorts and representations. Across these settings, pretrained gene representations broadly transferred mean expression but selectively improved spatial recovery, showing that cross-gene generalization in virtual spatial transcriptomics is not a single capability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chen, T., Hicks, S. C.. 2026-09-18. Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics. https://doi.org/10.64898/2026.09.15.751768
Cite the original work for its findings. Save a collection to share your selection of sources.