bioRxiv · 10.1101/2025.11.04.686583
Adaptive resampling for improved machine learning in imbalanced single-cell datasets
Abstract
While machine learning models trained on single-cell transcriptomics data have shown great promise in providing biological insights, existing tools struggle to effectively model underrepresented and out-of-distribution cellular features or states. We present a generalizable Adaptive Resampling (AR) approach that addresses these limitations and enhances single-cell representation learning by resampling data based on its learned latent structure in an online, adaptive manner concurrent with model training. Experiments on gene expression reconstruction, cell type classification, and perturbation response prediction tasks demonstrate that the proposed AR training approach leads to significantly improved downstream performance across datasets and metrics. Additionally, it enhances the quality of learned cellular embeddings compared to standard training methods. Our results suggest that AR may serve as a valuable technique for improving representation learning and predictive performance in single-cell transcriptomic models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Navidi, Z., Thoutam, A., Hughes, M., Raghavan, S., Winter, P. S., Crawford, L., Amini, A. P.. 2025-11-05. Adaptive resampling for improved machine learning in imbalanced single-cell datasets. https://doi.org/10.1101/2025.11.04.686583
Cite the original work for its findings. Save a collection to share your selection of sources.