Search bioRxivSearch

Biology subjects

Aerts, S.

Publications and source records attributed to Aerts, S..

6 recordsLinked to original sources

ChIP-seq meta-analysis yields high quality training sets for enhancer classification

Genome-wide prediction of enhancers depends on high-quality positive and negative training sets. The use of ChIP-seq peaks as positive training data can be problematic due to high degrees of indirectly bound regions, and often poor overlap between experimental conditions.\n\nHere we explore meta-analysis of ChIP-seq data to generate high-quality training data for enhancer modeling. Our method is based on rank aggregation and identifies a core set of directly bound regions per transcription factor, exploiting between five and twenty ChIP-seq data sets per factor. We applied this method to six different transcription factors, namely TP53, REST, SOX2, GRHL2, HIF1A and PPARG. Sequence analysis and modeling of recurrently bound enhancers yielded distinct enhancer features for the different factors, whereby binding sites of REST and TP53 are strongly determined by their motif; binding of GRHL2 and SOX2 is determined by nucleosome positioning; and binding of PPARG and HIF1A depends on other transcription factors. In conclusion, meta-analysis of ChIP-seq peaks, and centering on motifs, allowed discovering new properties of transcription factor binding.

bioinformatics

Bcl6 promotes neurogenic conversion through transcriptional repression of multiple self-renewal-promoting extrinsic pathways.

During neurogenesis, progenitors switch from self-renewal to differentiation through the interplay of intrinsic and extrinsic cues, but how these are integrated remains poorly understood. Here we combine whole genome transcriptional and epigenetic analyses with in vivo functional studies and show that Bcl6, a transcriptional repressor known to promote neurogenesis, acts as a key driver of the neurogenic transition through direct silencing of a selective repertoire of genes belonging to multiple extrinsic pathways promoting self-renewal, most strikingly the Wnt pathway. At the molecular level, Bcl6 acts through both generic and pathway-specific mechanisms. Our data identify a molecular logic by which a single cell-intrinsic factor ensures robustness of neural cell fate transition by decreasing responsiveness to the extrinsic pathways that favor self-renewal.

neuroscience

Cis-topic modelling of single cell epigenomes

Single-cell epigenomics provides new opportunities to decipher genomic regulatory programs from heterogeneous samples and dynamic processes. We present a probabilistic framework called cisTopic, to simultaneously discover \"cis-regulatory topics\" and stable cell states from sparse single-cell epigenomics data. After benchmarking cisTopic on single-cell ATAC-seq data, single-cell DNA methylation data, and semi-simulated single-cell ChIP-seq data, we use cisTopic to predict regulatory programs in the human brain and validate these by aligning them with co-expression networks derived from single-cell RNA-seq data. Next, we performed a time-series single-cell ATAC-seq experiment after SOX10 perturbations in melanoma cultures, where cisTopic revealed dynamic regulatory topics driven by SOX10 and AP-1. Finally, machine learning and enhancer modelling approaches allowed to predict cell type specific SOX10 and SOX9 binding sites based on topic specific co-regulatory motifs. cisTopic is available as an R/Bioconductor package at http://github.com/aertslab/cistopic.

bioinformatics

A single-cell catalogue of regulatory states in the ageing Drosophila brain

The diversity of cell types and regulatory states in the brain, and how these change during ageing, remains largely unknown. Here, we present a single-cell transcriptome catalogue of the entire adult Drosophila melanogaster brain sampled across its lifespan. Both neurons and glia age through a process of \"regulatory erosion\", characterized by a strong decline of RNA content, and accompanied by increasing transcriptional and chromatin noise. We identify more than 50 cell types by specific transcription factors and their downstream gene regulatory networks. In addition to neurotransmitter types and neuroblast lineages, we find a novel neuronal cell state driven by datilografo and prospero. This state relates to neuronal birth order, the metabolic profile, and the activity of a neuron. Our single-cell brain catalogue reveals extensive regulatory heterogeneity linked to ageing and brain function and will serve as a reference for future studies of genetic variation and disease mutations.

genomics

SCENIC: Single-Cell Regulatory Network Inference And Clustering

Single-cell RNA-seq allows building cell atlases of any given tissue and infer the dynamics of cellular state transitions during developmental or disease trajectories. Both the maintenance and transitions of cell states are encoded by regulatory programs in the genome sequence. However, this regulatory code has not yet been exploited to guide the identification of cellular states from single-cell RNA-seq data. Here we describe a computational resource, called SCENIC (Single Cell rEgulatory Network Inference and Clustering), for the simultaneous reconstruction of gene regulatory networks (GRNs) and the identification of stable cell states, using single-cell RNA-seq data. SCENIC outperforms existing approaches at the level of cell clustering and transcription factor identification. Importantly, we show that cell state identification based on GRNs is robust towards batch-effects and technical-biases. We applied SCENIC to a compendium of single-cell data from the mouse and human brain and demonstrate that the proper combinations of transcription factors, target genes, enhancers, and cell types can be identified. Moreover, we used SCENIC to map the cell state landscape in melanoma and identified a gene regulatory network underlying a proliferative melanoma state driven by MITF and STAT and a contrasting network controlling an invasive state governed by NFATC2 and NFIB. We further validated these predictions by showing that two transcription factors are predominantly expressed in early metastatic sentinel lymph nodes. In summary, SCENIC is the first method to analyze scRNA-seq data using a network-centric, rather than cell-centric approach. SCENIC is generic, easy to use, and flexible, and allows for the simultaneous tracing of genomic regulatory programs and the mapping of cellular identities emerging from these programs. Availability: SCENIC is available as an R workflow based on three new R/Bioconductor packages: GENIE3, RcisTarget and AUCell. As scalable alternative to GENIE3, we also provide GRNboost, paving the way towards the network analysis across millions of single cells.

bioinformatics

Identification Of cis-Regulatory Mutations Generating De Novo Edges In Personalized Cancer Gene Regulatory Networks

The identification of functional non-coding mutations is a key challenge in the field of genomics, where whole-genome re-sequencing can swiftly generate a set of all genomic variants in a sample, such as a tumor biopsy. The size of the human regulatory landscape places a challenge on finding recurrent cis-regulatory mutations across samples of the same cancer type. Therefore, powerful computational approaches are required to sift through the tens of thousands of non-coding variants, to identify potentially functional variants that have an impact on the gene expression profile of the sample. Here we introduce an integrative analysis pipeline, called - cisTarget, to filter, annotate and prioritize non-coding variants based on their putative effect on the underlying 'personal' gene regulatory network. We first validate -cisTarget by re-analyzing three cases of oncogenic non-coding mutations, namely the TAL1 and LMO1 enhancer mutations in T-ALL, and the TERT promoter mutation in melanoma. Next, we re-sequenced the full genome of ten cancer cell lines of six different cancer types, and used matched transcriptome data and motif discovery to infer master regulators for each sample. We identified candidate functional non-coding mutations that generate de novo binding sites for these master regulators, and that result in the up-regulation of nearby oncogenic drivers. We finally validated the predictions using tertiary data including matched epigenome data. Our approach is generally applicable to re-sequenced cancer genomes, or other genomes, when a disease- or sample-specific gene signature is available for network inference. -cisTarget is available from http://mucistarget.aertslab.org.

genomics