Search bioRxivSearch

Biology subjects

Kyono, Y.

Publications and source records attributed to Kyono, Y..

3 recordsLinked to original sources

Integrating enhancer RNA signatures with diverse omics data identifies characteristics of transcription initiation in pancreatic islets

Identifying the tissue-specific molecular signatures of active regulatory elements is critical to understand gene regulatory mechanisms. Here, we identify transcription start sites (TSS) using cap analysis of gene expression (CAGE) across 57 human pancreatic islet samples. We identify 9,954 reproducible CAGE tag clusters (TCs), ~20% of which are islet-specific and occur mostly distal to known gene TSSs. We integrated islet CAGE data with histone modification and chromatin accessibility profiles to identify epigenomic signatures of transcription initiation. Using a massively parallel reporter assay, we validate transcriptional enhancer activity (5% FDR) for 2,279 of 3,378 (~68%) tested islet CAGE elements. TCs within accessible enhancers show higher enrichment to overlap type 2 diabetes genome-wide association study (GWAS) signals than existing islet annotations, which emphasizes the utility of mapping CAGE profiles in disease-relevant tissue. This work provides a high-resolution map of transcriptional initiation in human pancreatic islets with utility for dissecting functional enhancers at GWAS loci.

genomics

Chromatin information content landscapes inform transcription factor and DNA interactions

Interactions between transcription factors (TFs) and chromatin are fundamental to genome organization and regulation and, ultimately, cell state. Here, we use information theory to measure signatures of TF-chromatin interactions encoded in the patterns of the accessible genome, which we call chromatin information enrichment (CIE). We calculate CIE for hundreds of TF motifs across human tissues and identify two classes: low and high CIE. The 10-20% of TF motifs with high CIE associate with higher protein-DNA residence time, including different binding sites subclasses of the same TF, increased nucleosome phasing, specific protein domains, and the genetic control of both gene expression and chromatin accessibility. These results show that variations in the information content of chromatin architecture reflect functional biological variation, with implications for cell state dynamics and memory.

genomics

STARRPeaker: Uniform processing and accurate identification of whole human STARR-seq active regions

BackgroundHigh-throughput reporter assays, such as self-transcribing active regulatory region sequencing (STARR-seq), allow for unbiased and quantitative assessment of enhancers at a genome-wide scale. Recent advances in STARR-seq technology have employed progressively more complex genomic libraries and increased sequencing depths, to assay larger sized regions, up to the entire human genome. These advances necessitate a reliable processing pipeline and peak-calling algorithm. ResultsMost STARR-seq studies have relied on chromatin immunoprecipitation sequencing (ChIP-seq) processing pipelines. However, there are key differences in STARR-seq versus ChIP-seq. First, STARR-seq uses transcribed RNA to measure the activity of an enhancer, making an accurate determination of the basal transcription rate important. Second, STARR-seq coverage is highly non-uniform, overdispersed, and often confounded by sequencing biases, such as GC content and mappability. Lastly, here, we observed a clear correlation between RNA thermodynamic stability and STARR-seq readout, suggesting that STARR-seq may be sensitive to RNA secondary structure and stability. Considering these findings, we developed a negative-binomial regression framework for uniformly processing STARR-seq data, called STARRPeaker. In support of this, we generated whole-genome STARR-seq data from the HepG2 and K562 human cell lines and applied STARRPeaker to call enhancers. ConclusionsWe show STARRPeaker can unbiasedly detect active enhancers from both captured and whole-genome STARR-seq data. Specifically, we report [~]33,000 and [~]20,000 candidate enhancers from HepG2 and K562, respectively. Moreover, we show that STARRPeaker outperforms other peak callers in terms of identifying known enhancers with fewer false positives. Overall, we demonstrate an optimized processing framework for STARR-seq experiments can identify putative enhancers while addressing potential confounders.

bioinformatics