Search bioRxiv⌕ Search

Biology subjects

Ong, X.

Publications and source records attributed to Ong, X..

2 recordsLinked to original sources

Integrative Epigenomic and High-Throughput Functional Enhancer Profiling Reveals Determinants of Enhancer Heterogeneity in Gastric Cancer

BackgroundEnhancers are distal cis-regulatory elements required for cell-specific gene expression and cell fate determination. In cancer, enhancer variation has been proposed as a major cause of inter-patient heterogeneity - however, most predicted enhancer regions remain to be functionally tested. ResultsAnalyzing 128 epigenomic histone modification profiles of primary GC samples, normal gastric tissues, and GC cell lines, we report a comprehensive catalog of 75,730 recurrent predicted enhancers, the majority of which are tumor-associated in vivo (>50,000) and associated with lower somatic mutation rates inferred by whole-genome sequencing. Applying Capture-based Self-Transcribing Active Regulatory Region sequencing (CapSTARR-seq) to the enhancer catalog, we observed significant correlations between CapSTARR-seq functional activity and H3K27ac/H3K4me1 levels. Super-enhancer regions exhibited increased CapSTARR-seq signals compared to regular enhancers even when decoupled from native chromatin contexture. We show that combining histone modification and CapSTARR-seq functional enhancer data improves the prediction of enhancer-promoter interactions and pinpointing of germline single nucleotide polymorphisms (SNPs), somatic copy number alterations (SCNAs), and trans-acting TFs involved in GC expression. Specifically, we identified cancer-relevant genes (e.g. ING1, ARL4C) whose expression between patients is influenced by enhancer differences in genomic copy number and germline SNPs, and HNF4 as a master trans-acting factor associated with GC enhancer heterogeneity. ConclusionsOur study indicates that combining histone modification and functional assay data may provide a more accurate metric to assess enhancer activity than either platform individually, and provides insights into the relative contribution of genetic (cis) and regulatory (trans) mechanisms to GC enhancer functional heterogeneity.

cancer biology↗

A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines

The human genome contains more than 200,000 gene isoforms. However, different isoforms can be highly similar, and with an average length of 1.5kb remain difficult to study with short read sequencing. To systematically evaluate the ability to study the transcriptome at a resolution of individual isoforms we profiled 5 human cell lines with short read cDNA sequencing and Nanopore long read direct RNA, amplification-free direct cDNA, PCR-cDNA sequencing. The long read protocols showed a high level of consistency, with amplification-free RNA and cDNA sequencing being most similar. While short and long reads generated comparable gene expression estimates, they differed substantially for individual isoforms. We find that increased read length improves read-to-transcript assignment, identifies interactions between alternative promoters and splicing, enables the discovery of novel transcripts from repetitive regions, facilitates the quantification of full-length fusion isoforms and enables the simultaneous profiling of m6A RNA modifications when RNA is sequenced directly. Our study demonstrates the advantage of long read RNA sequencing and provides a comprehensive resource that will enable the development and benchmarking of computational methods for profiling complex transcriptional events at isoform-level resolution.

genomics↗