Search bioRxivSearch

Biology subjects

Imrichova, H.

Publications and source records attributed to Imrichova, H..

3 recordsLinked to original sources

ChIP-seq meta-analysis yields high quality training sets for enhancer classification

Genome-wide prediction of enhancers depends on high-quality positive and negative training sets. The use of ChIP-seq peaks as positive training data can be problematic due to high degrees of indirectly bound regions, and often poor overlap between experimental conditions.\n\nHere we explore meta-analysis of ChIP-seq data to generate high-quality training data for enhancer modeling. Our method is based on rank aggregation and identifies a core set of directly bound regions per transcription factor, exploiting between five and twenty ChIP-seq data sets per factor. We applied this method to six different transcription factors, namely TP53, REST, SOX2, GRHL2, HIF1A and PPARG. Sequence analysis and modeling of recurrently bound enhancers yielded distinct enhancer features for the different factors, whereby binding sites of REST and TP53 are strongly determined by their motif; binding of GRHL2 and SOX2 is determined by nucleosome positioning; and binding of PPARG and HIF1A depends on other transcription factors. In conclusion, meta-analysis of ChIP-seq peaks, and centering on motifs, allowed discovering new properties of transcription factor binding.

bioinformatics

SCENIC: Single-Cell Regulatory Network Inference And Clustering

Single-cell RNA-seq allows building cell atlases of any given tissue and infer the dynamics of cellular state transitions during developmental or disease trajectories. Both the maintenance and transitions of cell states are encoded by regulatory programs in the genome sequence. However, this regulatory code has not yet been exploited to guide the identification of cellular states from single-cell RNA-seq data. Here we describe a computational resource, called SCENIC (Single Cell rEgulatory Network Inference and Clustering), for the simultaneous reconstruction of gene regulatory networks (GRNs) and the identification of stable cell states, using single-cell RNA-seq data. SCENIC outperforms existing approaches at the level of cell clustering and transcription factor identification. Importantly, we show that cell state identification based on GRNs is robust towards batch-effects and technical-biases. We applied SCENIC to a compendium of single-cell data from the mouse and human brain and demonstrate that the proper combinations of transcription factors, target genes, enhancers, and cell types can be identified. Moreover, we used SCENIC to map the cell state landscape in melanoma and identified a gene regulatory network underlying a proliferative melanoma state driven by MITF and STAT and a contrasting network controlling an invasive state governed by NFATC2 and NFIB. We further validated these predictions by showing that two transcription factors are predominantly expressed in early metastatic sentinel lymph nodes. In summary, SCENIC is the first method to analyze scRNA-seq data using a network-centric, rather than cell-centric approach. SCENIC is generic, easy to use, and flexible, and allows for the simultaneous tracing of genomic regulatory programs and the mapping of cellular identities emerging from these programs. Availability: SCENIC is available as an R workflow based on three new R/Bioconductor packages: GENIE3, RcisTarget and AUCell. As scalable alternative to GENIE3, we also provide GRNboost, paving the way towards the network analysis across millions of single cells.

bioinformatics

Identification Of cis-Regulatory Mutations Generating De Novo Edges In Personalized Cancer Gene Regulatory Networks

The identification of functional non-coding mutations is a key challenge in the field of genomics, where whole-genome re-sequencing can swiftly generate a set of all genomic variants in a sample, such as a tumor biopsy. The size of the human regulatory landscape places a challenge on finding recurrent cis-regulatory mutations across samples of the same cancer type. Therefore, powerful computational approaches are required to sift through the tens of thousands of non-coding variants, to identify potentially functional variants that have an impact on the gene expression profile of the sample. Here we introduce an integrative analysis pipeline, called - cisTarget, to filter, annotate and prioritize non-coding variants based on their putative effect on the underlying 'personal' gene regulatory network. We first validate -cisTarget by re-analyzing three cases of oncogenic non-coding mutations, namely the TAL1 and LMO1 enhancer mutations in T-ALL, and the TERT promoter mutation in melanoma. Next, we re-sequenced the full genome of ten cancer cell lines of six different cancer types, and used matched transcriptome data and motif discovery to infer master regulators for each sample. We identified candidate functional non-coding mutations that generate de novo binding sites for these master regulators, and that result in the up-regulation of nearby oncogenic drivers. We finally validated the predictions using tertiary data including matched epigenome data. Our approach is generally applicable to re-sequenced cancer genomes, or other genomes, when a disease- or sample-specific gene signature is available for network inference. -cisTarget is available from http://mucistarget.aertslab.org.

genomics