Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.09.03.748766

Atlas-scale single-cell analysis beyond in-memory paradigm with scAtlasPy

Abstract

Single-cell atlases are rapidly outgrowing the memory capacity of standard workstations, challenging the in-memory paradigm underlying mainstream computational ecosystems. Here, scAtlasPy decouples scale of atlas from memory capacity by leveraging the disk-resident computing. It enables full-resolution analysis of a 100-million-cell atlas with only 42.9 GB peak memory, whereas state-of-the-art platforms are limited at 3 million cells with 512 GB memory. scAtlasPy achieves 137,745 cells/s, 10.4x faster than scDataset with 82.6% lower memory usage for random minibatch retrieval. Its extensible architecture offers a flexible platform for diverse atlas-scale analytical tasks, facilitating the discovery of complex cellular heterogeneity and functions in massive cell atlases.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xu, H., Ye, Y., Zhang, S., Xie, R., Li, J., Lin, J., Hu, Y., Gao, L.. 2026-09-08. Atlas-scale single-cell analysis beyond in-memory paradigm with scAtlasPy. https://doi.org/10.64898/2026.09.03.748766

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Physical priors improve performance of structure-based binding affinity models

Structure-based drug discovery is a widely used paradigm for the rational design of novel small molecule therapeutics. However, the benefits conferred by the use of structural information has seen limited adoption in machine learning, where ligand-only ("2D") models are still the industry standard for molecular property or binding affinity prediction. Structure-based ("3D") ML models for binding-affinity prediction promise to present a clear advantage, but have not yet overtaken existing 2D models. Here, we show that physics-based priors can improve predictive performance of structure-based models by comparing different model architectures with varying physical priors on several prediction tasks. We present the Modular Training and Evaluation of Neural Networks (mtenn) package, where we decompose affinity prediction into separate steps of embedding structure into learned representations and combining those embeddings into a predicted binding affinity. We consider both E(3)-invariant and E(3)-equivariant architectures to determine the importance of encoding roto-translational inductive biases, as well as different methods for combining learned embeddings. By first optimizing several aspects of model construction using the general purpose PDBBind dataset, we are able to improve the performance and data efficiency of structure-based models. When subsequently trained and evaluated on the COVID Moonshot small molecule drug discovery dataset, our tuned models perform on par with industry standard ligand-only models. Our decomposed model framework highlights that encoding some physical priors improves model performance, while more complex biases such as equivariance offer limited benefit. Additionally, structure-based models generalize better to an unseen target and display higher training efficiency. Overall, these results emphasize that structure-based models benefit from their ability to incorporate physics-informed constraints, giving promising directions for model architecture development. These results also suggest that the strength of these models may be in tasks specifically aimed at generalizability, providing guidelines for their use in early-stage drug discovery campaigns.

bioinformatics↗

Rapid Shift Toward Pulsed Field Ablation and Precision Risk Stratification in High-Impact Atrial Fibrillation Research

Conventional bibliometrics rely on lifetime citations, obscuring immediate shifts in cardiovascular research paradigms. To track emerging trends in atrial fibrillation management, we performed a comparative bibliometric analysis of the 50 highest-cited original research articles per year in OpenAlex topic T10065 across consecutive 2023 (Class of 2025) and 2024 (Class of 2026) publication cohorts. Articles were ranked using a fixed 18-month post-publication citation window, and extracted concepts were normalized into 719 canonical topics and 32 parent themes using large language model curation. Concept frequency tracking demonstrated a swift technological shift, with pulsed field ablation showing the largest topic frequency increase (+0.08, from 0.30 to 0.38) to become the leading canonical topic in 2024, displacing conventional thermal pulmonary vein isolation (-0.18, 0.40 to 0.22). Simultaneously, stroke prevention focus shifted toward refined predictive modeling, with increases in Risk Stratification and Predictive Models (+0.08, 0.30 to 0.38) and CHA2DS2-VASc scoring (+0.08, from 0.10 to 0.18). High-impact atrial fibrillation research is rapidly pivoting toward non-thermal ablation safety profiling and precision risk stratification, highlighting the utility of fixed-window concept mining for capturing real-time scientific evolution. Online explorer of the result is available at https://pri.pepkio.com.

bioinformatics↗

Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models

Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

bioinformatics↗