Search bioRxivSearch

Biology subjects

Jain, M. S.

Publications and source records attributed to Jain, M. S..

3 recordsLinked to original sources

MultiMAP: Dimensionality Reduction and Integration of Multimodal Data

Multimodal data is rapidly growing in many fields of science and engineering, including single-cell biology. We introduce MultiMAP, an approach for dimensionality reduction and integration of multiple datasets. MultiMAP recovers a single manifold on which all of the data resides and then projects the data into a single low-dimensional space so as to preserve the structure of the manifold. It is based on a framework of Riemannian geometry and algebraic topology, and generalizes the popular UMAP algorithm1 to the multimodal setting. MultiMAP can be used for visualization of multimodal data, and as an integration approach that enables joint analyses. MultiMAP has several advantages over existing integration strategies for single-cell data, including that MultiMAP can integrate any number of datasets, leverages features that are not present in all datasets (i.e. datasets can be of different dimensionalities), is not restricted to a linear mapping, can control the influence of each dataset on the embedding, and is extremely scalable to large datasets. We apply MultiMAP to the integration of a variety of single-cell transcriptomics, chromatin accessibility, methylation, and spatial data, and show that it outperforms current approaches in preservation of high-dimensional structure, alignment of datasets, visual separation of clusters, transfer learning, and runtime. On a newly generated single-cell Assay for Transposase-Accessible Chromatin using sequencing (scATAC-seq) and single-cell RNA-seq (scRNA-seq) dataset of the human thymus, we use MultiMAP to integrate cells along a temporal trajectory. This enables the quantitative comparison of transcription factor expression and binding site accessibility over the course of T cell differentiation, revealing patterns of transcription factor kinetics.

bioinformatics

Comprehensive mapping of tissue cell architecture via integrated single cell and spatial transcriptomics

The spatial organization of cell types in tissues fundamentally shapes cellular interactions and function, but the high-throughput spatial mapping of complex tissues remains a challenge. We present [c]ell2location, a principled and versatile Bayesian model that integrates single-cell and spatial transcriptomics to map cell types in situ in a comprehensive manner. We show that [c]ell2location outperforms existing tools in accuracy and comprehensiveness and we demonstrate its utility by mapping two complex tissues. In the mouse brain, we use a new paired single nucleus and spatial RNA-sequencing dataset to map dozens of cell types and identify tissue regions in an automated manner. We discover novel regional astrocyte subtypes including fine subpopulations in the thalamus and hypothalamus. In the human lymph node, we resolve spatially interlaced immune cell states and identify co-located groups of cells underlying tissue organisation. We spatially map a rare pre-germinal centre B-cell population and predict putative cellular interactions relevant to the interferon response. Collectively our results demonstrate how [c]ell2location can serve as a versatile first-line analysis tool to map tissue architectures in a high-throughput manner.

genomics

Predicting tumour mutational burden from histopathological images using multiscale deep learning

Tumour mutational burden (TMB) is an important biomarker for predicting response to immunotherapy in cancer patients. Gold-standard measurement of TMB is performed using whole exome sequencing (WES), which is not available at most hospitals owing to its high cost, operational complexity, and long turnover times. We developed a machine learning algorithm, Image2TMB, which can predict TMB from readily available lung adenocarcinoma histopathological images. Image2TMB integrates the predictions of three deep learning models that operate at different resolution scales (5X, 10X, and 20X magnification) to determine if the TMB of a cancer is high or low. On a held-out set of patients, Image2TMB achieves an area under the precision recall curve of 0.92, an average precision of 0.89, and has the predictive power of a targeted sequencing panel of approximately 100 genes. This study demonstrates that it is possible to infer genomic features from histopathology images, and potentially opens avenues for exploring genotype-phenotype relationships.

bioinformatics