Search bioRxiv⌕ Search

Biology subjects

Pandit, V.

Publications and source records attributed to Pandit, V..

2 recordsLinked to original sources

SpatialPrompt: spatially aware scalable and accurate tool for spot deconvolution and clustering in spatial transcriptomics

Spatial transcriptomics has advanced our understanding of tissue biology by enabling sequencing while preserving spatial coordinates. In sequencing-based spatial technologies, each measured spot typically consists of multiple cells. Deconvolution algorithms are required to decipher the cell-type distribution at each spot. Existing spot deconvolution algorithms for spatial transcriptomics often neglect spatial coordinates and lack scalability as datasets get larger. We introduce SpatialPrompt, a spatially aware and scalable method for spot deconvolution as well as domain identification for spatial transcriptomics. Our method integrates gene expression, spatial location, and single-cell RNA sequencing (scRNA-seq) reference data to infer cell-type proportions of spatial spots accurately. At the core, SpatialPrompt uses non-negative ridge regression and an iterative approach inspired by graph neural network (GNN) to capture the local microenvironment information in the spatial data. Quantitative assessments on the human prefrontal cortex dataset demonstrated the superior performance of our tool for spot deconvolution and domain identification. Additionally, SpatialPrompt accurately decipher the spatial niches of the mouse cortex and the hippocampus regions that are generated from different protocols. Furthermore, consistent spot deconvolution prediction from multiple references on the mouse kidney spatial dataset showed the impressive robustness of the tool. In response to this, SpatialPromptDB database is developed to provide compatible scRNA-seq references with cell-type annotations for seamless integration. In terms of scalability, SpatialPrompt is the only method performing spot deconvolution and clustering in less than 2 minutes for large spatial datasets with 50,000 spots. SpatialPrompt tool along with the SpatialPromptDB database are publicly available as open source software for large-scale spatial transcriptomics analysis.

bioinformatics↗

Comparison of Dimensionality Reduction and Clustering Methods for Single-Cell Transcriptomics Data

Dimensionality reduction (DR) methods are applied to extract relevant features from inherently high dimensional and noisy single-cell RNA sequencing (scRNA-seq) data. Choice of DR method could influence the performance of clustering algorithm and subsequent analysis outcomes. We performed a benchmarking study of seven popular DR methods and four clustering algorithms widely used for scRNA-seq datasets. For this purpose, we used three publicly available real scRNA-seq datasets. The performance was evaluated using two clustering metrics viz. adjusted random index (ARI) and normalized mutual index (NMI). We also compared our results with a similar study published by Xiang and colleagues. Overall, we observed higher ARI and NMI scores for DR methods when compared with Xiangs study. We also noticed several differences between our and Xiangs study. Noteworthy, three methods, namely, Independent Component Analysis (ICA), t-Distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP) performed consistently well across three datasets. Linear method ICA was best performer on Segerstolpe dataset, while nonlinear methods UMAP and t-SNE best performed on Deng and Chu datasets, respectively. Neural network-based methods Variational Autoencoder (VAE) and Deep Count Autoencoder (DCA) could not perform well probably due to their sensitivity to hyperparameters and overfitting. Among clustering methods, Gaussian Mixture Models (GMMs) performed consistently well across datasets. This might be because GMMs are the universal approximators of posterior probability densities. We conclude that performance of different DR methods is more dataset dependent and for various scRNA-seq datasets different algorithms are more suited and there is no one-fit-all method.

bioinformatics↗