Search bioRxiv⌕ Search

Biology subjects

Grabski, I. N.

Publications and source records attributed to Grabski, I. N..

4 recordsLinked to original sources

Bayesian Multi-Study Non-Negative Matrix Factorization for Mutational Signatures

AO_SCPLOWBSTRACTC_SCPLOWMutational signatures shed insight into the range of mutational processes giving rise to tumors and allow a better understanding of cancer origin. They are typically identified from high-throughput sequencing data of cancer genomes using non-negative matrix factorization (NMF), and many such techniques have been developed towards this aim. However, it is often of particular interest to compare mutational signatures across multiple conditions, e.g. to understand which signatures are present across different treatments, or to identify signatures that are shared or specific across cancer types. Existing techniques within the NMF context only allow decomposition within a single dataset, so that integrating results across multiple conditions requires running separate analyses on each dataset, followed by subjective and manual comparisons of the identified signatures. To address this issue, we propose a Bayesian multi-study NMF method that jointly decomposes multiple studies or conditions to identify signatures that are common, specific, or partially shared by any subset. We propose two models: a "discovery-only" model that estimates de novo signatures in a completely unsupervised manner, and a "recovery-discovery" model that builds informative priors from previously known signatures to both update the estimates of these signatures and identify any novel signatures. We then further extend these models to estimate the effects of sample-level covariates on the exposures to each signature, enforcing sparsity through a non-local spike-and-slab prior. We demonstrate our approach on a range of simulations, and apply our method to colorectal cancer samples to show its utility.

bioinformatics↗

Significance Analysis for Clustering with Single-Cell RNA-Sequencing Data

AO_SCPLOWBSTRACTC_SCPLOWUnsupervised clustering of single-cell RNA-sequencing data enables the identification and discovery of distinct cell populations. However, the most widely used clustering algorithms are heuristic and do not formally account for statistical uncertainty. Many popular pipelines use clustering stability methods to assess the algorithms output and decide on the number of clusters. However, we find that by not addressing known sources of variability in a statistically rigorous manner, these analyses lead to overconfidence in the discovery of novel cell-types. We extend a previous method for Gaussian data, Significance of Hierarchical Clustering (SHC), to propose a model-based hypothesis testing approach that incorporates significance analysis into the clustering algorithm and permits statistical evaluation of clusters as distinct cell populations. We also adapt this approach to permit statistical assessment on the clusters reported by any algorithm. We benchmarked our approach on real-world datasets against popular clustering workflows, demonstrating improved performance. To show its practical utility, we applied it to the Human Lung Cell Atlas and an atlas of the mouse cerebellar cortex. We identified several cases of over-clustering, leading to false discoveries, as well as under-clustering, resulting in the failure to identify new subpopulations that our method was able to detect.

bioinformatics↗

Differentially methylated regions and methylation QTLs for teen depression and early puberty in the Fragile Families Child Wellbeing Study

AO_SCPLOWBSTRACTC_SCPLOWThe Fragile Families Child Wellbeing Study (FFCWS) is a longitudinal cohort of ethnically diverse and primarily low socioeconomic status children and their families in the U.S. Here, we analyze DNA methylation data collected from 748 FFCWS participants in two waves of this study, corresponding to participant ages 9 and 15. Our primary goal is to leverage the DNA methylation data from these two time points to study methylation associated with two key traits in adolescent health that are over-represented in these data: Early puberty and teen depression. We first identify differentially methylated regions (DMRs) for depression and early puberty. We then identify DMRs for the interaction effects between these two conditions and age by including interaction terms in our regression models to understand how age-related changes in methylation are influenced by depression or early puberty. Next, we identify methylation quantitative trait loci (meQTLs) using genotype data from the participants. We also identify meQTLs with epistatic effects with depression and early puberty. We find enrichment of our interaction meQTLs with functional categories of the genome that contribute to the heritability of co-morbid complex diseases. We replicate our meQTLs in data from the GoDMC study. This work leverages the important focus of the FFCWS data on disadvantaged children to shed light on the methylation states associated with teen depression and early puberty, and on how genetic regulation of methylation is affected in adolescents with these two conditions.

genomics↗

Probabilistic gene expression signatures identify cell-types from single cell RNA-seq data

AO_SCPLOWBSTRACTC_SCPLOWSingle-cell RNA sequencing (scRNA-seq) quantifies gene expression for individual cells in a sample, which allows distinct cell-type populations to be identified and characterized. An important step in many scRNA-seq analysis pipelines is the annotation of cells into known cell-types. While this can be achieved using experimental techniques, such as fluorescence-activated cell sorting, these approaches are impractical for large numbers of cells. This motivates the development of data-driven cell-type annotation methods. We find limitations with current approaches due to the reliance on known marker genes or from overfitting because of systematic differences between studies or batch effects. Here, we present a statistical approach that leverages public datasets to combine information across thousands of genes, uses a latent variable model to define cell-type-specific barcodes and account for batch effect variation, and probabilistically annotates cell-type identity. The barcoding approach also provides a new way to discover marker genes. Using a range of datasets, including those generated to represent imperfect real-world reference data, we demonstrate that our approach substantially outperforms current reference-based methods, in particular when predicting across studies. Our approach also demonstrates that current approaches based on unsupervised clustering lead to false discoveries related to novel cell-types.

genomics↗