bioRxiv · 10.1101/2024.12.09.627458
Tissue-Specific Cell Type Annotation with Supervised Representation Learning using Split Vector Quantization and Its Comparisons with Single-cell Foundation Models
Abstract
Cell-type annotation in single-cell data involves identifying and labeling the cell types based on their gene expression profiles or molecular features. Recently, with advances in single-cell foundation models (FMs), unsupervised annotation and transfer learning with FMs have been explored for cell-type annotation tasks. However, because FMs are usually pre-trained in an unsupervised manner on data spanning a wide variety of tissues and cell types, their representations for specific tissues may lack specificity and become overly generalized. In this work, we propose a novel supervised representation learning method using split-vector-quantization, single-cell Vector-Quantization Classifier (scVQC). We evaluated scVQC against both supervised and unsupervised representation learning approaches, with a focus on foundation models pretrained on large-scale single-cell datasets, such as scBERT and scGPT. The experimental results highlight the importance of label supervision in cell-type annotation tasks and demonstrate that the learned codebook effectively profiles and distinguishes different cell types.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Heryanto, Y. D., Zhang, Y.-z., Imoto, S.. 2024-12-13. Tissue-Specific Cell Type Annotation with Supervised Representation Learning using Split Vector Quantization and Its Comparisons with Single-cell Foundation Models. https://doi.org/10.1101/2024.12.09.627458
Cite the original work for its findings. Save a collection to share your selection of sources.