Search bioRxiv⌕ Search

Biology subjects

Szalata, A.

Publications and source records attributed to Szalata, A..

4 recordsLinked to original sources

CellFlow enables generative single-cell phenotype modeling with flow matching

High-content phenotypic screens provide a powerful strategy for studying biological systems, but the scale of possible perturbations and cell states makes exhaustive experiments unfeasible. Computational models that are trained on existing data and extrapolate to correctly predict outcomes in unseen contexts have the potential to accelerate biological discovery. Here, we present CellFlow, a flexible framework based on flow matching that can model single cell phenotypes induced by complex perturbations. We apply CellFlow to various phenotypic screens, accurately predicting expression responses to a wide range of perturbations, including cytokine stimulation, drug treatments and gene knockouts. CellFlow successfully modeled developmental perturbations at the whole-embryo scale and guided cell fate and organoid engineering by predicting heterogeneous cell populations arising from combinatorial morphogen treatments and by performing a virtual organoid protocol screen. Taken together, CellFlow has the potential to accelerate discovery from phenotypic screens by learning from existing data and generating phenotypes induced by unseen conditions.

bioinformatics↗

Multimodal weakly supervised learning to identify disease-specific changes in single-cell atlases

To deliver clinically relevant insights from large patient cohorts profiled with single-cell technologies, a key challenge is to relate sample-level and single-cell measurements. We present MultiMIL, a deep learning framework that applies attention-based multiple-instance learning for phenotype prediction and cell state identification. We applied MultiMIL to peripheral blood mononuclear cells from COVID-19 patients, the Human Lung Cell Atlas, and a spatial proteomics breast cancer dataset, demonstrating how our model can be utilized to find phenotype-associated cell states, learn phenotype-informed sample representations, and expand disease signatures.

bioinformatics↗

An integrated transcriptomic cell atlas of human neural organoids

Neural tissues generated from human pluripotent stem cells in vitro (known as neural organoids) are becoming useful tools to study human brain development, evolution and disease. The characterization of neural organoids using single-cell genomic methods has revealed a large diversity of neural cell types with molecular signatures similar to those observed in primary human brain tissue. However, it is unclear which domains of the human nervous system are covered by existing protocols. It is also difficult to quantitatively assess variation between protocols and the specific cell states in organoids as compared to primary counterparts. Single-cell transcriptome data from primary tissue and neural organoids derived with guided or un-guided approaches and under diverse conditions combined with large-scale integrative analyses make it now possible to address these challenges. Recent advances in computational methodology enable the generation of integrated atlases across many data sets. Here, we integrated 36 single-cell transcriptomics data sets spanning 26 protocols into one integrated human neural organoid cell atlas (HNOCA) totaling over 1.7 million cells. We harmonize cell type annotations by incorporating reference data sets from the developing human brain. By mapping to the developing human brain reference, we reveal which primary cell states have been generated in vitro, and which are under-represented. We further compare transcriptomic profiles of neuronal populations in organoids to their counterparts in the developing human brain. To support rapid organoid phenotyping and quantitative assessment of new protocols, we provide a programmatic interface to browse the atlas and query new data sets, and showcase the power of the atlas to annotate new query data sets and evaluate new organoid protocols. Taken together, the HNOCA will be useful to assess the fidelity of organoids, characterize perturbed and diseased states and facilitate protocol development in the future.

developmental biology↗

Biologically relevant integration of transcriptomics profiles from cancer cell lines, patient-derived xenografts and clinical tumors using deep neural networks

Cell lines and patient-derived xenografts are essential to cancer research, however, the results derived from such models often lack clinical translatability, as these models do not fully recapitulate the complex cancer biology. It is critically important to better understand the systematic differences between cell lines, xenografts and clinical tumors, and to be able to identify pre-clinical models that sufficiently resemble the biological characteristics of clinical tumors across different cancers. On another side, direct comparison of transcriptional profiles from pre-clinical models and clinical tumors is infeasible due to the mixture of technical artifacts and inherent biological signals. To address these challenges, we developed MOBER, Multi-Origin Batch Effect Remover method, to simultaneously extract biologically meaningful embeddings and remove batch effects from transcriptomic datasets of different origin. MOBER consists of two neural networks: conditional variational autoencoder and source discriminator neural network that is trained in adversarial fashion. We applied MOBER on transcriptional profiles from 932 cancer cell lines, 434 patient-derived tumor xenografts and 11159 clinical tumors and identified pre-clinical models with greatest transcriptional fidelity to clinical tumors, and models that are transcriptionally unrepresentative of their respective clinical tumors. MOBER can conserve the biological signals from the original datasets, while generating embeddings that do not encode confounder information. In addition, it allows for transformation of transcriptional profiles of pre-clinical models to resemble the ones of clinical tumors, and therefore can be used to improve the clinical translation of insights gained from pre-clinical models. As a batch effect removal method, MOBER can be applied widely to transcriptomics datasets of different origin, allowing for integration of multiple datasets simultaneously.

bioinformatics↗