bioRxiv · 10.64898/2026.02.11.705441
Improving SCVI for low-count cells through self-supervised augmentation
Abstract
When analyzing single-cell RNA sequencing data with SCVI, low-UMI cells typically need to be filtered, because their learned representations lack meaningful biological signal. We show that this is caused by a specific mechanism: as UMI depth decreases, the SCVI encoder maps cells towards a learned bias point, collapsing their representations regardless of cell identity. This phenomenon is distinct from classical posterior collapse driven by KL regularization. By augmenting training with binomial thinning and adding a cross-correlation loss between original and thinned cell embeddings, the encoder learns representations that preserve cell type identity, experimental condition differences, and sample-level variation at lower UMI depths, without sacrificing reconstruction quality. These modifications extend the range of usable cells, enabling analysis of cells that would typically be discarded due to low molecule counts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Svensson, V.. 2026-02-13. Improving SCVI for low-count cells through self-supervised augmentation. https://doi.org/10.64898/2026.02.11.705441
Cite the original work for its findings. Save a collection to share your selection of sources.