Search bioRxiv⌕ Search

Biology subjects

Hrovatin, K.

Publications and source records attributed to Hrovatin, K..

2 recordsLinked to original sources

Biologically informed deep learning to infer gene program activity in single cells

The increasing availability of large-scale single-cell datasets has enabled the detailed description of cell states across multiple biological conditions and perturbations. In parallel, recent advances in unsupervised machine learning, particularly in transfer learning, have enabled fast and scalable mapping of these new single-cell datasets onto reference atlases. The resulting large-scale machine learning models however often have millions of parameters, rendering interpretation of the newly mapped datasets challenging. Here, we propose expiMap, a deep learning model that enables interpretable reference mapping using biologically understandable entities, such as curated sets of genes and gene programs. The key concept is the substitution of the uninterpretable nodes in an autoencoders bottleneck by labeled nodes mapping to interpretable lists of genes, such as gene ontologies, biological pathways, or curated gene sets, for which activities are learned as constraints during reconstruction. This is enabled by the incorporation of predefined gene programs into the reference model, and at the same time allowing the model to learn de novo new programs and refine existing programs during reference mapping. We show that the model retains similar integration performance as existing methods while providing a biologically interpretable framework for understanding cellular behavior. We demonstrate the capabilities of expiMap by applying it to 15 datasets encompassing five different tissues and species. The interpretable nature of the mapping revealed unreported associations between interferon signaling via the RIG-I/MDA5 and GPCRs pathways, with differential behavior in CD8+ T cells and CD14+ monocytes in severe COVID-19, as well as the role of annexins in the cellular communications between lymphoid and myeloid compartments for explaining patient response to the applied drugs. Finally, expiMap enabled the direct comparison of a diverse set of pancreatic beta cells from multiple studies where we observed a strong, previously unreported correlation between the unfolded protein response and asparagine N-linked glycosylation. Altogether, expiMap enables the interpretable mapping of single cell transcriptome data sets across cohorts, disease states and other perturbations.

bioinformatics↗

Transcriptional milestones in Dictyostelium development

Development of the social amoeba Dictyostelium discoideum begins by starvation of single cells and ends in multicellular fruiting bodies 24 hours later. These major morphological changes are accompanied by sweeping gene expression changes, encompassing nearly half of the 13,000 genes in the genome. To explore the relationships between the transcriptome and developmental morphogenesis, we performed time-series RNA-sequencing analysis of the wild type and 20 mutant strains with altered morphogenesis. These strains exhibit arrest at different developmental stages, accelerated development, or terminal morphologies that are not typically seen in the wild type. Considering eight major morphological transitions, we identified 1,371 milestone genes whose expression changes sharply between two consecutive transitions. We also identified 1,099 genes as members of 21 regulons, which are groups of genes that remain coordinately regulated despite the genetic, temporal, and developmental perturbations in the dataset. The gene annotations in these milestones and regulons validate known transitions and reveal several new physiological and functional transitions during development. For example, we found that DNA replication genes are co-regulated with cell division genes, so they are co-expressed in mid-development even though chromosomal DNA is not replicated at that time. Altogether, the dataset includes 486 transcriptional profiles, across developmental and genetic conditions, that can be used to identify new relationships between gene expression and developmental processes and to improve gene annotations. We demonstrate the utility of this resource by showing that the cycles of aggregation and disaggregation observed in allorecognition-defective mutants involve a dedifferentiation process. We also show unexpected variability and sensitivity to genetic background and developmental conditions in two commonly used genes, act6 and act15, and robustness of the coaA gene. Finally, we propose that gpdA should be used as a standard for mRNA quantitation because it is less sensitive to genetic background and developmental conditions than commonly used standards. The dataset is available for democratized exploration without the need for programming skills through the web application dictyExpress and the data mining environment Orange.

genomics↗