Search bioRxivSearch

Biology subjects

Sengupta, D.

Publications and source records attributed to Sengupta, D..

7 recordsLinked to original sources

deepMc: deep Matrix Completion for imputation of single cell RNA-seq data

Single cell RNA-seq has fueled discovery and innovation in medicine over the past few years and is useful for studying cellular responses at individual cell resolution. But, due to paucity of starting RNA, the data acquired is highly sparse. To address this, We propose a deep matrix factorization based method, deepMc, to impute missing values in gene-expression data. For the deep architecture of our approach, We draw our motivation from great success of deep learning in solving various Machine learning problems. In this work, We support our method with positive results on several evaluation metrics like clustering of cell populations, differential expression analysis and cell type separability.

bioinformatics

ROSeq: A rank based approach to modelling gene expression in single cells

1Systematic delineation of complex biological systems is an ever-challenging and resource-intensive process. Single cell transcriptomics allows us to study cell-to-cell variability in complex tissues at an unprecedented resolution. Accurate modeling of gene expression plays a critical role in the statistical determination of tissue-specific gene expression patterns. In the past few years, considerable efforts have been made to identify appropriate parametric models for single cell expression data. The zero-inflated version of Poisson/Negative Binomial and Log-Normal distributions have emerged as the most popular alternatives due to their ability to accommodate high dropout rates, as commonly observed in single cell data. While the majority of the parametric approaches directly model expression estimates, we explore the potential of modeling expression-ranks, as robust surrogates for transcript abundance. Here we examined the performance of the Discrete Generalized Beta Distribution (DGBD) on real data and devised a Wald-type test for comparing gene expression across two phenotypically divergent groups of single cells. We performed a comprehensive assessment of the proposed method, to understand its advantages as compared to some of the existing best practice approaches. Besides striking a reasonable balance between Type 1 and Type 2 errors, we concluded that ROSeq, the proposed differential expression test is exceptionally robust to expression noise and scales rapidly with increasing sample size. For wider dissemination and adoption of the method, we created an R package called ROSeq, and made it available on the Bioconductor platform.

genomics

McImpute: Matrix completion based imputation for single cell RNA-seq data

MotivationSingle cell RNA sequencing has been proved to be revolutionary for its potential of zooming into complex biological systems. Genome wide expression analysis at single cell resolution, provides a window into dynamics of cellular phenotypes. This facilitates characterization of transcriptional heterogeneity in normal and diseased tissues under various conditions. It also sheds light on development or emergence of specific cell populations and phenotypes. However, owing to the paucity of input RNA, a typical single cell RNA sequencing data features a high number of dropout events where transcripts fail to get amplified.\n\nResultsWe introduce mcImpute, a low-rank matrix completion based technique to impute dropouts in single cell expression data. On a number of real datasets, application of mcImpute yields significant improvements in separation of true zeros from dropouts, cell-clustering, differential expression analysis, cell type separability, performance of dimensionality reduction techniques for cell visualization and gene distribution.\n\nAvailability and Implementationhttps://github.com/aanchalMongia/McImpute_scRNAseq

bioinformatics

Disproportionate feedback interactions govern cell-type specific proliferation in mammalian cells

In mammalian cells, the critical decision to maintain quiescence over proliferation commitment in and around the G1-S transition depends on more than one intertwined feedback interaction, and is highly cell-type dependent. However, the precise role played by these individual feedback regulations, in order to generate such diverse nature of proliferation commitment, is still poorly understood. Herein, we propose a generic mathematical model of G1-S transition in mammalian cells that not only reconciles distinct single cell experimental observations in a cell-type specific manner, but also makes experimentally testable non-intuitive predictions. Intriguingly, The model analysis reveals that the feedback motifs responsible for the G1-S transition act in a disparate fashion to organize the cell-type specific proliferation response in different mammalian cells. The proposed model, in principle, can be effectively tuned to explore proliferation dynamics in a cell-type specific way to gain crucial insights about novel therapeutic intervention to prevent unwanted cellular proliferation.

systems biology

Subtle alteration in microRNA dynamics accounts for differential nature of cellular proliferation

In the G1 phase of the mammalian cell cycle, a bi-stable steady state dynamics of the transcription factor E2F ensures that only a certain threshold level of the growth factor can induce a high expression level (on state) of E2F to initiate either normal or abnormal cellular proliferation or even apoptosis. A group of microRNAs known as the mir-17-92 cluster, which specifically inhibits E2F, can simultaneously influence the threshold level of growth factor required for E2F activation, and the corresponding expression level of E2F in the on state. However, mir-17-92 cluster can function as either oncogene or tumor suppressor in a cell-type specific manner for reasons that still remain illusive. Here we put forward a deterministic mathematical model for Myc/E2F/mir-17-92 network that demonstrates how the experimentally observed mir-17-92 mediated differential nature of the cellular proliferation can be reconciled by having conflicting steady state dynamics of E2F for different cell types. While a 2-D bifurcation study of the model rationalizes the reason behind the contrasting E2F dynamics, an intuitive sensitivity analysis of the model parameters predicts that by exclusively altering the mir-17-92 related part of the network, it is possible to experimentally manipulate the cellular proliferation in a cell-type specific fashion for therapeutic intervention.

systems biology

dropClust: Efficient clustering of ultra-large scRNA-seq data

Droplet based single cell transcriptomics has recently enabled parallel screening of tens of thousands of single cells. Clustering methods that scale for such high dimensional data without compromising accuracy are scarce. We exploit Locality Sensitive Hashing, an approximate nearest neighbor search technique to develop a de novo clustering algorithm for large-scale single cell data. On a number of real datasets, dropClust outperformed the existing best practice methods in terms of execution time, clustering accuracy and detectability of minor cell sub-types.

genomics

FORKS: Finding Orderings Robustly using K-means and Steiner trees

Recent advances in single cell RNA-seq technologies have provided researchers with unprecedented details of transcriptomic variation across individual cells. However, it has not been straightforward to infer differentiation trajectories from such data, due to the parameter-sensitivity of existing methods. Here, we present Finding Orderings Robustly using k-means and Steiner trees (FORKS), an algorithm that pseudo-temporally orders cells and thereby infers bifurcating state trajectories. FORKS, which is a generic method, can be applied to both single-cell and bulk differentiation data. It is a semi-supervised approach, in that it requires the user to specify the starting point of the time course. We systematically benchmarked FORKS and eight other pseudo-time estimation algorithms on six benchmark datasets, and found it to be more accurate, more reproducible, and more memory-efficient than existing methods for pseudo-temporal ordering. Another major advantage of our approach is its robustness - FORKS can be used with default parameter settings on a wide range of datasets.

bioinformatics