Search bioRxivSearch

Biology subjects

Colantuoni, C.

Publications and source records attributed to Colantuoni, C..

5 recordsLinked to original sources

Decomposing cell identity for transfer learning across cellular measurements, platforms, tissues, and species.

New approaches are urgently needed to glean biological insights from the vast amounts of single cell RNA sequencing (scRNA-Seq) data now being generated. To this end, we propose that cell identity should map to a reduced set of factors which will describe both exclusive and shared biology of individual cells, and that the dimensions which contain these factors reflect biologically meaningful relationships across different platforms, tissues and species. To find a robust set of dependent factors in large-scale scRNA- Seq data, we developed a Bayesian non-negative matrix factorization (NMF) algorithm, scCoGAPS. Application of scCoGAPS to scRNA-Seq data obtained over the course of mouse retinal development identified gene expression signatures for factors associated with specific cell types and continuous biological processes. To test whether these signatures are shared across diverse cellular contexts, we developed projectR to map biologically disparate datasets into the factors learned by scCoGAPS. Because projecting these dimensions preserve relative distances between samples, biologically meaningful relationships/factors will stratify new data consistent with their underlying processes, allowing labels or information from one dataset to be used for annotation of the other--a machine learning concept called transfer learning. Using projectR, data from multiple datasets was used to annotate latent spaces and reveal novel parallels between developmental programs in other tissues, species and cellular assays. Using this approach we are able to transfer cell type and state designations across datasets to rapidly annotate cellular features in a new dataset without a priori knowledge of their type, identify a species-specific signature of microglial cells, and identify a previously undescribed subpopulation of neurosecretory cells within the lung. Together, these algorithms define biologically meaningful dimensions of cellular identity, state, and trajectories that persist across technologies, molecular features, and species.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=174 HEIGHT=200 SRC=\"FIGDIR/small/395004_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (81K):\norg.highwire.dtl.DTLVardef@dd1c07org.highwire.dtl.DTLVardef@5b1109org.highwire.dtl.DTLVardef@bb6714org.highwire.dtl.DTLVardef@16c66f0_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics

Genome-scale transcriptional regulatory network models of psychiatric and neurodegenerative disorders

Genetic and genomic studies suggest an important role for transcriptional regulatory changes in brain diseases, but roles for specific transcription factors (TFs) remain poorly understood. We integrated human brain-specific DNase I footprinting and TF-gene co-expression to reconstruct a transcriptional regulatory network (TRN) model for the human brain, predicting the brain-specific binding sites and target genes for 741 TFs. We used this model to predict core TFs involved in psychiatric and neurodegenerative diseases. Our results suggest that disease-related transcriptomic and genetic changes converge on small sets of disease-specific regulators, with distinct networks underlying neurodegenerative vs. psychiatric diseases. Core TFs were frequently implicated in a disease through multiple mechanisms, including differential expression of their target genes, disruption of their binding sites by disease-associated SNPs, and associations of the genetic loci encoding these TFs with disease risk. We validated our models predictions through systematic comparison to publicly available ChIP-seq and TF perturbation studies and through experimental studies in primary human neural stem cells. Combined genetic and transcriptional evidence supports roles for neuronal and microglia-enriched, MEF2C-regulated networks in Alzheimers disease; an oligodendrocyte-enriched, SREBF1-regulated network in schizophrenia; and a neural stem cell and astrocyte-enriched, POU3F2-regulated network in bipolar disorder. We provide our models of brain-specific TF binding sites and target genes as a resource for network analysis of brain diseases.

genetics

Placental gene expression mediates the interaction between obstetrical history and genetic risk for schizophrenia

Defining the environmental context in which genes enhance susceptibility can provide insight into the pathogenesis of complex disorders, like schizophrenia. Here we show that the intrauterine and perinatal environment modulates the association of schizophrenia with genomic risk, as measured with polygenic risk scores (PRS) based primarily on GWAS significant variants. Genomic risk interacts with intrauterine and perinatal complications (Early Life Complications, ELCs) in each of three independent samples from USA, Italy and Germany (overall n= 1693, p= 6e-05). In each sample, the liability of schizophrenia explained by PRS is nominally more than five times greater in the presence of a history of ELCs compared with its absence. In each sample, patients with positive ELC histories have higher PRS than patients without ELCs, which is further confirmed in two additional patient samples from Germany and Japan (overall n=2038, p= 1e-04). The gene set based on the schizophrenia loci interacting with ELCs is highly expressed in multiple placental compartments and dynamically regulated in placenta from complicated in comparison with normal pregnancies. The same genes are differentially up-regulated in placentae from male compared with female offspring. The interaction between genomic risk and ELCs is mainly driven by GWAS significant loci enriched for genes highly expressed in the various placenta samples. Molecular pathway analyses based on the genes not driving this interaction reflect previous analyses about schizophrenia risk-genes, while genes highly and differentially expressed in placentae implicate an orthogonal biology involving cellular stress. These results suggest that the most significant genetic variants detected by current schizophrenia GWAS contribute to risk in part by converging on a developmental trajectory sensitive to events affecting placentation, which may underlie the male preponderance of schizophrenia and offer new insights into primary prevention.

genetics

Developmental And Genetic Regulation Of The Human Cortex Transcriptome In Schizophrenia

GWAS have identified 108 loci that confer risk for schizophrenia, but risk mechanisms for individual loci are largely unknown. Using developmental, genetic, and illness-based RNA sequencing expression analysis, we characterized the human brain transcriptome around these loci and found enrichment for developmentally regulated genes with novel examples of shifting isoform usage across pre- and post-natal life. We found widespread expression quantitative trait loci (eQTLs), including many with transcript specificity and previously unannotated sequence that were independently replicated. We leveraged this eQTL database to show that 48.1% of risk variants for schizophrenia associated with nearby expression. Within patients and controls, we implemented a novel algorithm for RNA quality adjustment, and identified 237 genes significantly associated with diagnosis that replicated in an independent case-control dataset. These genes implicated synaptic processes and were strongly regulated in early development (p < 10-20). These data offer new targets for modeling schizophrenia risk in cellular systems.

neuroscience

PatternMarkers and Genome-Wide CoGAPS Analysis in Parallel Sets (GWCoGAPS) for data-driven detection of novel biomarkers via whole transcriptome Non-negative matrix factorization (NMF)

SummaryNon-negative Matrix Factorization (NMF) algorithms associate gene expression with biological processes (e.g., time-course dynamics or disease subtypes). Compared with univariate associations, the relative weights of NMF solutions can obscure biomarkers. Therefore, we developed a novel PatternMarkers statistic to extract genes for biological validation and enhanced visualization of NMF results. Finding novel and unbiased gene markers with PatternMarkers requires whole-genome data. However, NMF algorithms typically do not converge for the tens of thousands of genes in genome-wide profiling. Therefore, we also developed Genome-Wide CoGAPS Analysis in Parallel Sets (GWCoGAPS), the first robust whole genome Bayesian NMF using the sparse, MCMC algorithm, CoGAPS. This software contains analytic and visualization tools including a Shiny web application, patternMatcher, which are generalized for any NMF. Using these tools, we find granular brain-region and cell-type specific signatures with corresponding biomarkers in GTex data, illustrating GWCoGAPS and patternMarkers ascertainment of data-driven biomarkers from whole-genome data.\n\nAvailabilityPatternMarkers & GWCoGAPS are in the CoGAPS Bioconductor package (3.5) under the GPL license.\n\nContactgsteinobrien@jhmi.edu; ccolantu@jhmi.edu; ejfertig@jhmi.edu

bioinformatics