Search bioRxiv⌕ Search

Biology subjects

Mowlaei, M. E.

Publications and source records attributed to Mowlaei, M. E..

3 recordsLinked to original sources

De novo transformer modeling improves recovery of genetic cell types from sparse single-cell RNA sequencing

Single-cell RNA sequencing (scRNA-seq) simultaneously provides gene-expression profiles and genetic variants from individual cells, creating an opportunity to relate cellular phenotypes to their somatic evolutionary histories. However, delineation of genetic type (GTs) from scRNA-seq remains difficult because most variant positions are unobserved in individual cells and the observed base calls contain substantial false-positive and false-negative errors. We evaluated some existing phylogenetic and imputation methods using one simulated dataset and two empirical tumor datasets. We found that extreme sparsity prevented reliable recovery of known or independently inferred GTs when multiple GTs were present. This led us to adapt the STICI transformer architecture to train a separate model de novo on each sparse cell-variant (CV) matrix. These data-specific models predicted millions of missing bases, greatly reducing matrix sparsity. Phylogenetic analyses of the imputed CV matrices showed substantially improved recovery of GTs in both simulated and empirical datasets. In the empirical dataset, transformer-based analysis also suggested finer-scale genetic structure within some previously reported GTs that was not apparent with the existing methods. These results demonstrate that highly sparse scRNA-seq datasets contain substantially more recoverable lineage information than previously appreciated and that de novo transformer modeling provides an effective approach for recovering much of this hidden information. Nevertheless, sequencing errors persisted, limiting reconstruction of cellular lineage structure and leaving significant room for methodological improvement before expression phenotypes can be examined reliably in the context of their cellular evolutionary relationships.

bioinformatics↗

TRUHiC: A TRansformer-embedded U-2 Net to enhance Hi-C data for 3D chromatin structure characterization

High-throughput chromosome conformation capture sequencing (Hi-C) is a key technology for studying the three-dimensional (3D) structure of genomes and chromatin folding. Hi-C data reveals underlying patterns of genome organization, such as topologically associating domains (TADs) and chromatin loops, with critical roles in transcriptional regulation and disease etiology and progression. However, the sparsity of existing Hi-C data often hinders robust and reliable inference of 3D structures. Hence, we propose TRUHiC, a new computational method that leverages recent state-of-the-art deep generative modeling to augment low-resolution Hi-C data for the characterization of 3D chromatin structures. By applying TRUHiC to real low-resolution Hi-C data from the GM12329 cell line and across other publicly available Hi-C data for human and mice, we demonstrate that the augmented data significantly improve the characterization of TADs and loops across diverse cell lines and species. We further present a pre-trained TRUHiC on human lymphoblastoid cell lines that can be adaptable and transferable to improve chromatin characterization of various cell lines, tissues, and species.

bioinformatics↗

Split-Transformer Impute (STI): Genotype Imputation Using a Transformer-Based Model

MotivationDespite recent advances in sequencing technologies, genome-scale datasets continue to have missing bases and genomic segments. Such incomplete datasets can undermine downstream analyses, such as disease risk prediction and association studies. Consequently, the imputation of missing information is a common pre-processing step for which many methodologies have been developed. However, the imputation of genotypes of certain genomic regions and variants, including large structural variants, remains a challenging problem. ResultsHere, we present a transformer-based deep learning framework, called a split-transformer impute (STI) model, for accurate genome-scale genotype imputation. Empowered by the attention-based transformer model, STI can be trained for any collection of genomes automatically using self-supervision. STI handles multi-allelic genotypes naturally, unlike other models that need special treatments. STI models automatically learned genome-wide patterns of linkage disequilibrium (LD), evidenced by much higher imputation accuracy in high LD regions. Also, STI models trained through sporadic masking for self-supervision performed well in imputing systematically missing information. Our imputation results on the human 1000 Genomes Project show that STI can achieve high imputation accuracy, comparable to the state-of-the-art genotype imputation methods, with the additional capability to impute multi-allelic structural variants and other types of genetic variants. Moreover, STI showed excellent performance without needing any special presuppositions about the patterns in the underlying data when applied to a collection of yeast genomes, pointing to easy adaptability and application of STI to impute missing genotypes in any species.

genomics↗