Search bioRxiv⌕ Search

Biology subjects

Moinfar, A. A.

Publications and source records attributed to Moinfar, A. A..

5 recordsLinked to original sources

Unsupervised Deep Disentangled Representation of Single-Cell Omics

Deep generative models have become central to single-cell omics analysis, but their latent spaces remain difficult to interpret biologically. Linear factor models offer dimension-wise interpretability, but often lack the nonlinear flexibility, scalability, and integration quality required for large multi-batch atlases. We bridge this gap with Disentangled Representation Variational Inference (DRVI), an unsupervised deep generative model that learns dimension-wise interpretable representations for single-cell omics without supervised priors on cell types or biological processes. DRVI achieves this through additive decoder subnetworks combined with log-sum-exp pooling, enabling disentangled nonlinear gene programs. Across atlases, perturbation screens, and developmental datasets, DRVI separates cell identity, signaling pathways, stress responses, developmental trajectories, and perturbation effects into distinct interpretable factors. These factors identify rare migratory dendritic cells and recover coherent combinatorial perturbation programs in CRISPR screens. Systematic benchmarks show that this interpretability does not reduce integration performance. DRVI recovers biological factors more accurately while maintaining competitive integration quality. Altogether, DRVI provides a practical route to nonlinear single-cell modeling with factor-level interpretability.

bioinformatics↗

Pertpy: an end-to-end framework for perturbation analysis

Advances in single-cell technology have enabled the measurement of cell-resolved molecular states across a variety of cell lines and tissues under a plethora of genetic, chemical, environmental, or disease perturbations. Current methods focus on differential comparison or are specific to a particular task in a multi-condition setting with purely statistical perspectives. The quickly growing number, size, and complexity of such studies requires a scalable analysis framework that takes existing biological context into account. Here, we present pertpy, a Python-based modular framework for the analysis of large-scale perturbation single-cell experiments. Pertpy provides access to harmonized perturbation datasets and metadata databases along with numerous fast and user-friendly implementations of both established and novel methods such as automatic metadata annotation or perturbation distances to efficiently analyze perturbation data. As part of the scverse ecosystem, pertpy interoperates with existing libraries for the analysis of single-cell data and is designed to be easily extended.

bioinformatics↗

Multimodal weakly supervised learning to identify disease-specific changes in single-cell atlases

To deliver clinically relevant insights from large patient cohorts profiled with single-cell technologies, a key challenge is to relate sample-level and single-cell measurements. We present MultiMIL, a deep learning framework that applies attention-based multiple-instance learning for phenotype prediction and cell state identification. We applied MultiMIL to peripheral blood mononuclear cells from COVID-19 patients, the Human Lung Cell Atlas, and a spatial proteomics breast cancer dataset, demonstrating how our model can be utilized to find phenotype-associated cell states, learn phenotype-informed sample representations, and expand disease signatures.

bioinformatics↗

Integrating single-cell RNA-seq datasets with substantial batch effects

Integration of single-cell RNA-sequencing (scRNA-seq) datasets has become a standard part of the analysis, with conditional variational autoencoders (cVAE) being among the most popular approaches. Increasingly, researchers are asking to map cells across challenging cases such as cross-organs, species, or organoids and primary tissue, as well as different scRNA-seq protocols, including single-cell and single-nuclei. Current computational methods struggle to harmonize datasets with such substantial differences, driven by technical or biological variation. Here, we propose to address these challenges for the popular cVAE-based approaches by introducing and comparing a series of regularization constraints. The two commonly used strategies for increasing batch correction in cVAEs, that is Kullback-Leibler divergence (KL) regularization strength tuning and adversarial learning, suffer from substantial loss of biological information. Therefore, we adapt, implement, and assess alternative regularization strategies for cVAEs and investigate how they improve batch effect removal or better preserve biological variation, enabling us to propose an optimal cVAE-based integration strategy for complex systems. We show that using a VampPrior instead of the commonly used Gaussian prior not only improves the preservation of biological variation but also unexpectedly batch correction. Moreover, we show that our implementation of cycle-consistency loss leads to significantly better biological preservation than adversarial learning implemented in the previously proposed GLUE model. Additionally, we do not recommend relying only on the KL regularization strength tuning for increasing batch correction, as it removes both biological and batch information without discriminating between the two. Based on our findings, we propose a new model that combines VampPrior and cycle-consistency loss. We show that using it for datasets with substantial batch effects improves downstream interpretation of cell states and biological conditions. To ease the use of the newly proposed model, we make it available in the scvi-tools package as an external model named sysVI. Moreover, in the future, these regularization techniques could be added to other established cVAE-based models to improve the integration of datasets with substantial batch effects.

bioinformatics↗

Population-level integration of single-cell datasets enables multi-scale analysis across samples

The increasing generation of population-level single-cell atlases with hundreds or thousands of samples has the potential to link demographic and technical metadata with high-resolution cellular and tissue data in homeostasis and disease. Constructing such comprehensive references requires large-scale integration of heterogeneous cohorts with varying metadata capturing demographic and technical information. Here, we present single-cell population level integration (scPoli), a semi-supervised conditional deep generative model for data integration, label transfer and query-to-reference mapping. Unlike other models, scPoli learns both sample and cell representations, is aware of cell-type annotations and can integrate and annotate newly generated query datasets while providing an uncertainty mechanism to identify unknown populations. We extensively evaluated the method and showed its advantages over existing approaches. We applied scPoli to two population-level atlases of lung and peripheral blood mononuclear cells (PBMCs), the latter consisting of roughly 8 million cells across 2,375 samples. We demonstrate that scPoli allows atlas-level integration and automatic reference mapping with label transfer. It can explain sample-level biological and technical variations such as disease, anatomical location and assay by means of its novel sample embeddings. We use these embeddings to explore sample-level metadata, enable automatic sample classification and guide a data integration workflow. scPoli also enables simultaneous sample-level and cell-level analysis of gene expression patterns, revealing genes associated with batch effects and the main axes of between-sample variation. We envision scPoli becoming an important tool for population-level single-cell data integration facilitating atlas use but also interpretation by means of multi-scale analyses.

genomics↗