Search bioRxiv⌕ Search

Biology subjects

Ganapathi, T.

Publications and source records attributed to Ganapathi, T..

2 recordsLinked to original sources

Species-specific small models for cell type classification approach the performance of large single cell foundation models

Accurate cross-species cell type classification remains a key evaluation task in single-cell transcriptomics. Recent foundation models trained on millions of single-cell profiles demonstrate great in-distribution and out-of-distribution performance on this task, but their large parameter counts and substantial computational costs limit accessibility and interpretability. Here, we introduce CytoType, a simple and interpretable model for cell type classification that leverages pre-trained ESM-2 protein embeddings of protein-coding transcripts. By learning linear, cell-type-specific weights over transcript embeddings, without relying on gene count information, CytoType achieves F1 scores comparable to or exceeding those of large-scale transformer-based models. We further developed ESM-Cell Embedding (ESM-CE), an even simpler variant that only averages ESM-2 embeddings across expressed genes, which also performs competitively against foundation models. Both CytoType and ESM-CE are trained on species-specific data, maintaining high accuracy when classifying cell types with orders of magnitude fewer parameters compared to larger foundation models. For example, for human tissues, the average performance gap between CytoType and the best foundation model was 0.053 F1 points while CytoType uses 10,000x fewer trainable parameters. Additionally, we quantified the contribution of ESM-2 embeddings to cell type classification tasks and demonstrated a three fold reduction in the performance gap between CytoType and the best foundation model for nine species. Finally, we show that CytoTypes learned gene weights are biologically interpretable.

cell biology↗

VariantFormer: A hierarchical transformer integrating DNA sequences with genetic variations and regulatory landscapes for personalized gene expression prediction

1Accurately predicting gene expression from DNA sequence remains a central challenge in human genetics. Current sequence-based models overlook natural genetic variation across individuals, while population-based models are restricted to variants observed within specific cohorts. Here, we present VariantFormer, a 1.2-billion-parameter transformer that predicts gene-level RNA abundance directly from personalized diploid genomes. Trained on 21,004 genome-transcriptome pairs from 2,330 donors, VariantFormer achieves state-of-the-art performance across both sequence- and population-based prediction tasks, while generalizing better to out-of-distribution contexts--including somatic mutation settings in cancer cell lines--and main-taining robustness across ancestries. Beyond expression prediction, VariantFormer improves eQTL effect size estimation compared to prior methods, with notable gains for lower-frequency and ancestry-specific variants. In applications to Alzheimers disease, VariantFormer gene embeddings prioritize likely causal genes and relevant tissue contexts, and in silico mutagenesis of known APOE alleles faithfully recovers known risk modifying effects. Together, these results establish VariantFormer as a scalable, diploid-aware framework for variant interpretation and personalized gene expression modeling across tissues and populations.

genomics↗