Search bioRxiv⌕ Search

Biology subjects

Yabuta, Y.

Publications and source records attributed to Yabuta, Y..

2 recordsLinked to original sources

scEGOT: Single-cell trajectory inference framework based on entropic Gaussian mixture optimal transport

Time-series scRNA-seq data have opened a door to elucidate cell differentiation, and in this context, the optimal transport theory has been attracting much attention. However, there remain critical issues in interpretability and computational cost. We present scEGOT, a comprehensive framework for single-cell trajectory inference, as a generative model with high interpretability and low computational cost. Applied to the human primordial germ cell-like cell (PGCLC) induction system, scEGOT identified the PGCLC progenitor population and bifurcation time of segregation. Our analysis shows TFAP2A is insufficient for identifying PGCLC progenitors, requiring NKX1-2. Additionally, MESP1 and GATA6 are also crucial for PGCLC/somatic cell segregation. These findings shed light on the mechanism that segregates PGCLC from somatic lineages. Notably, not limited to scRNA-seq, scEGOTs versatility can extend to general single-cell data like scATAC-seq, and hence has the potential to revolutionize our understanding of such datasets and, thereby also, developmental biology.

bioinformatics↗

Resolution of the curse of dimensionality in single-cell RNA sequencing data analysis

Single-cell RNA sequencing (scRNA-seq) can determine gene expression in numerous individual cells simultaneously, promoting progress in the biomedical sciences. However, scRNA-seq data are high-dimensional with substantial technical noise, including dropouts. During analysis of scRNA-seq data, such noise engenders a statistical problem known as the curse of dimensionality (COD). Based on high-dimensional statistics, we herein formulate a noise reduction method, RECODE (resolution of the curse of dimensionality), for high-dimensional data with random sampling noise. We show that RECODE consistently eliminates COD in relevant scRNA-seq data with unique molecular identifiers. RECODE does not involve dimension reduction and recovers expression values for all genes, including lowly expressed genes, realizing precise delineation of cell-fate transitions and identification of rare cells with all gene information. Compared to other representative imputation methods, RECODE employs different principles and exhibits superior overall performance in cell-clustering and single-cell level analysis. The RECODE algorithm is parameter-free, data-driven, deterministic, and high-speed, and notably, its applicability can be predicted based on the variance normalization performance. We propose RECODE as a general strategy for preprocessing noisy high-dimensional data.

bioinformatics↗