Search bioRxiv⌕ Search

Biology subjects

Vroland, C.

Publications and source records attributed to Vroland, C..

4 recordsLinked to original sources

The human RNA-DNA interactome is cell type-specific and dynamic

More than twenty years ago, the FANTOM consortium uncovered that mammalian genomes are pervasively transcribed, revealing multitudes of RNAs with unknown functions. A subset of these transcripts has since then been linked to transcriptional control and to chromatin organization via their ability to interact with DNA, suggesting that chromatin-associated RNAs could be key players in genome regulation. Although recent technological advances now enable the mapping of genome-wide RNA-DNA contacts, a lack of analyses integrating these methods with other genomic features and across multiple cellular contexts hinders our comprehensive understanding of the principles underlying RNA-DNA interactions and of their biological importance. As part of the FANTOM6 project, we thus generated RNA-DNA interaction maps in 16 different human cell types, then combined these contacts with multiple layers of other genomic data to investigate how patterns of interaction between RNA and DNA relate to chromatin organization and function. We show that the RNA-DNA interactome is highly dynamic yet reproducibly organized in cell-type specific networks, constituted of a great diversity of interactions that vary in function of their distance, the nature of their sources and the chromatin state of their targets. In particular, we detected numerous regulatory elements that exhibit marked changes in activity when differentially bound by transcripts, implying that thousands of RNA-DNA interactions can play a mechanistic role in gene expression. This regulatory function correlates with RNA-protein interactions and significantly associates with cell type-relevant and disorder-related traits. In addition to providing essential resources for future research in RNA-mediated chromatin regulation, cellular biology and human diseases, our study thus establishes the RNA-DNA interactome as a new genome regulatory layer that defines and maintains cellular identity and behavior.

genomics↗

Controlling for DNA dosage with Whole Genome Sequencing improves ATAC-seq peak calling

The Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) is a scalable and sensitive method for profiling chromatin accessibility, enabling the identification of cis-regulatory elements (CREs) that govern gene expression in diverse cellular contexts. Although ATAC-seq is routinely applied to both bulk and single-cell samples, we reveal that its peak calling process is compromised by local biases in DNA dosage, arising not only from copy number variations (CNVs) but also from DNA replication timing (RT). These biases can distort read coverage and compromise peak detection accuracy. As part of the FANTOM consortiums efforts to elucidate genomic regulation, we propose enhancing the MACS pipeline by integrating whole-genome sequencing (WGS) data to account for local DNA dosage effects, analogous to the use of input controls in ChIP-seq analyses. By incorporating WGS data, we demonstrate an increase in both the number and width of ATAC peaks, with improved proximity to transcription start sites (TSSs). WGS-controlled ATAC peaks exhibit canonical CRE epigenetic marks and are enriched for trait- and disease-associated genetic variants. Furthermore, the number of WGS-controlled peaks correlates more strongly with gene expression levels compared to peaks called without WGS control. Collectively, these results demonstrate that integrating WGS as a control significantly enhances the accuracy of ATAC-seq peak calling. Critically, we show that even low-depth WGS data is sufficient to improve peak calling performance, making this approach both cost-effective and readily adoptable for routine analyses. To ensure accessibility and reproducibility, we implemented this method as an open-source Nextflow pipeline. By challenging the assumption of uniform genomic visibility, our approach also holds broad implications for other DNA sequencing-based technologies.

bioinformatics↗

A genome-wide, machine learning-guided exploration of the cis-regulatory code involved in neuronal differentiation

Gene expression is controlled by proximal and distal cis-regulatory elements (CREs), containing DNA motifs bound by various transcription factors (TFs). Other sequence features, such as specific k-mers or low complexity regions, have also been implicated [1-3]. However, in a dynamic biological process such as cell differentiation, we lack an understanding of how the transcriptional activity of CREs progressively change and what sequence features underlie these transitions, which may reflect common and/or coordinated regulatory processes. Here, we use single-cell ATAC-seq and RNA-seq to follow, at a genome scale, CREs along differentiation of induced pluripotent stem cells into cortical neurons and develop a method to automatically identify the diversity of CRE profiles and their underlying sequence features. We propose a machine-learning guided clustering algorithm, STOIC (Statistical learning TO Inform Clustering), that jointly learns an unsupervised clustering of the CREs in the space of the activity profiles and a supervised predictor associated with each cluster in the DNA-sequence space. STOIC is specifically designed to provide readily interpretable results. We show that the method identifies CRE profiles associated with highly predictive sequence features and outperforms methods solely concerned with co-activity clustering on this task. Orthogonal data collected in the same settings link the inferred CRE clusters to specific enhancer or promoter signatures. Furthermore, we show that the DNA features unveiled by STOIC reflect biologically relevant regulators and offer a valuable basis to dissect elements of the cisregulatory grammar. Finally, we demonstrate the general applicability of STOIC by analyzing five bulk CAGE datasets of human cells responding to various treatments.

bioinformatics↗

Evidence of distal regulations orchestrated by RNAs initiating at short tandem repeats

Short Tandem Repeats (STRs), also called microsatellites, correspond to tandemly repeated short DNA motifs (1 to 6 bp) and are one of the most polymorphic and abundant repetitive elements in the human genome. Variations of their length (i.e. number of consecutive repeats) have been implicated in gene expression regulation (termed expression(e)STRs). Using captrapping followed by long read sequencing, we discovered that STRs can host transcription start sites, the presence of which depends mainly on STR flanking sequences. Here, we investigate the effect of SNPs located in these sequences and ask whether STR-initiating RNAs have regulatory potential. First, we develop fully interpretable deep learning-based models, called Modular Neural Networks, able to predict, for each STR class, the level of RNAs using 101bp-long sequences encompassing STRs. Analysis of MNN filters allows us to identify multiple regulatory elements and candidate transcription factors. Second, leveraging genome sequencing and gene expression data from the Genotype-Tissue Expression project, we use the output of our models to link the levels of STR-initiating RNAs to the expression of nearby genes. We identify 14,340 significant associations (coined RNA(r)STRs) and illustrate how this novel resource can help interpret non-coding variants associated with complex traits and diseases. Third, we unveil an intricate transcriptional interplay between STR-initiating RNAs and Alu repeats that may couple their regulatory actions, extending both the nature and the functional importance of non-coding transcription and shedding new light on the complexity of distal regulations orchestrated by repeated sequences.

genomics↗