Search bioRxivSearch

Biology subjects

Marsico, A.

Publications and source records attributed to Marsico, A..

4 recordsLinked to original sources

Identifying lncRNA-mediated regulatory modules via ChIA-PET network analysis

BackgroundAlthough several studies have provided insights into the role of long non-coding RNAs (lncRNAs), the majority of them has unknown function. Recent evidence has shown the importance of both lncR-NAs and chromatin interactions in transcriptional regulation. Although network-based methods, mainly exploiting gene-lncRNA co-expression, have been applied to characterize lncRNA of unknown function by means of guilt-by-association strategies, no method exists which combines co-expression analysis with 3D chromatin interaction data.\n\nResultsTo better understand the function of chromatin interactions in the context of lncRNA-mediated gene regulation, we have developed a multi-step graph analysis approach to examine the RNA polymerase II ChIA-PET chromatin interaction network in the K562 human cell line. We have annotated the network with gene and lncRNA coordinates, and chromatin states from the ENCODE project. We used centrality measures, as well as an adaptation of our previously developed Markov State Models (MSM) clustering method, to gain a better understanding of lncRNAs in transcriptional regulation. The novelty of our approach resides into the detection of fuzzy regulatory modules based on network properties and their optimization based on co-expression analysis between genes and gene-lncRNA pairs. This results in our method returning more bona fide regulatory modules than other state-of-the art approaches for clustering on graphs.\n\nConclusionsInterestingly, we find that lncRNA network hubs tend to be significantly enriched in disease association, positional conservation and enhancer-like functions. We validated regulatory functions for well known lncRNAs, such as MALAT1 and the enhancer-like lncRNA FALEC. In addition, by investigating the modular structure of bigger components we show that we can propose regulatory functional mechanisms for uncharacterized lncRNAs, such FLJ37453, RP11442N24 B.1 and LINC00910.

systems biology

pysster: Learning Sequence and Structure Motifs in DNA and RNA Sequences using Convolutional Neural Networks

SummaryConvolutional neural networks (CNNs) have been shown to perform exceptionally well in a variety of tasks, including biological sequence classification. Available implementations, however, are usually optimized for a particular task and difficult to reuse. To enable researchers to utilize these networks more easily we implemented pysster, a Python package for training CNNs on biological sequence data. Sequences are classified by learning sequence and structure motifs and the package offers an automated hyper-parameter optimization procedure and options to visualize learned motifs along with information about their positional and class enrichment. The package runs seamlessly on CPU and GPU and provides a simple interface to train and evaluate a network with a handful lines of code. Using an RNA A-to-I editing data set and CLIP-seq binding site sequences we demonstrate that pysster classifies sequences with higher accuracy than other methods and is able to recover known sequence and structure motifs.\n\nAvailabilitypysster is freely available at https://github.com/budach/pysster.\n\nContactbudach@molgen.mpg.de, marsico@molgen.mpg.de

bioinformatics

Chromatin-release of the long ncRNA A-ROD is required for transcriptional activation of its target gene DKK1

Long non-coding RNAs (ncRNAs) are involved in both positive and negative regulation of transcription. Long ncRNAs are often enriched in the nucleus and at chromatin but whether chromatin-release plays a functional role is unknown. Here, we used epigenetic marks, expression level and strength of chromatin interactions to group long ncRNAs and find that those engaged in strong chromatin interactions are less enriched at chromatin in MCF-7 cells, suggesting a functional involvement of chromatin-release of long ncRNAs in transcriptional regulation. To study this further, we identify the long ncRNA A-ROD, an activating regulator of the Wnt signaling inhibitor DKK1. We show that A-ROD enhances transcription elongation of DKK1 in an RNA-dependent manner and that A-ROD recruits EBP1 to the DKK1 promoter. Our data suggest that the activating function depends on the release of A-ROD from chromatin, and further identify a functional regulatory interaction mediated by A-ROD in the transcription activation of DKK1. We propose that the release of a subset of long ncRNAs is important for their function, adding a new mechanistic perspective to the subcellular localization of long ncRNAs.

molecular biology

PureCLIP: Capturing target-specific protein-RNA interaction footprints from single-nucleotide CLIP-seq data

iCLIP and eCLIP techniques facilitate the detection of protein-RNA interaction sites at high resolution, based on diagnostic events at crosslink sites. However, previous methods do not explicitly model the specifics of iCLIP and eCLIP truncation patterns and possible biases. We developed PureCLIP, a hidden Markov model based approach, which simultaneously performs peak calling and individual crosslink site detection. It explicitly incorporates RNA abundances and, for the first time, non-specific sequence biases. On both simulated and real data, PureCLIP is more accurate in calling crosslink sites than other state-of-the-art methods and has a higher agreement across replicates. Link: https://github.com/skrakau/PureCLIP.

bioinformatics