Search bioRxivSearch

Biology subjects

Hongyu Zhao

Publications and source records attributed to Hongyu Zhao.

8 recordsLinked to original sources

Systematic tissue-specific functional annotation of the human genome highlights immune-related DNA elements for late-onset Alzheimer’s disease

Continuing efforts from large international consortia have made genome-wide epigenomic and transcriptomic annotation data publicly available for a variety of cell and tissue types. However, synthesis of these datasets into effective summary metrics to characterize the functional non-coding genome remains a challenge. Here, we present GenoSkyline-Plus, an extension of our previous work through integration of an expanded set of epigenomic and transcriptomic annotations to produce high-resolution, single tissue annotations. After validating our annotations with a catalog of tissue-specific non-coding elements previously identified in the literature, we apply our method using data from 127 different cell and tissue types to present an atlas of heritability enrichment across 45 different GWAS traits. We show that broader organ system categories (e.g. immune system) increase statistical power in identifying biologically relevant tissue types for complex diseases while annotations of individual cell types (e.g. monocytes or B-cells) provide deeper insights into disease etiology. Additionally, we use our GenoSkyline-Plus annotations in an in-depth case study of late-onset Alzheimers disease (LOAD). Our analyses suggest a strong connection between LOAD heritability and genetic variants contained in regions of the genome functional in monocytes. Furthermore, we show that LOAD shares a similar localization of SNPs to monocyte-functional regions with Parkinsons disease. Overall, we demonstrate that integrated genome annotations at the single tissue level provide a valuable tool for understanding the etiology of complex human diseases. Our GenoSkyline-Plus annotations are freely available at http://genocanyon.med.yale.edu/GenoSkyline.\n\nAuthor SummaryAfter years of community efforts, many experimental and computational approaches have been developed and applied for functional annotation of the human genome, yet proper annotation still remains challenging, especially in non-coding regions. As complex disease research rapidly advances, increasing evidence suggests that non-coding regulatory DNA elements may be the primary regions harboring risk variants in human complex diseases. In this paper, we introduce GenoSkyline-Plus, a principled annotation framework to identify tissue and cell type-specific functional regions in the human genome through integration of diverse high-throughput epigenomic and transcriptomic data. Through validation of known non-coding tissue-specific regulatory regions, enrichment analyses on 45 complex traits, and an in-depth case study of neurodegenerative diseases, we demonstrate the ability of GenoSkyline-Plus to accurately identify tissue-specific functionality in the human genome and provide unbiased, genome-wide insights into the genetic basis of human complex diseases.

Genetics

Leveraging Functional Annotations in Genetic Risk Prediction for Human Complex Diseases

Genome wide association studies have identified numerous regions in the genome associated with hundreds of human diseases. Building accurate genetic risk prediction models from these data will have great impacts on disease prevention and treatment strategies. However, prediction accuracy remains moderate for most diseases, which is largely due to the challenges in identifying all the disease-associated variants and accurately estimating their effect sizes. We introduce AnnoPred, a principled framework that incorporates diverse functional annotation data to improve risk prediction accuracy, and demonstrate its performance on multiple human complex diseases.

Bioinformatics

Three distinct velocities of elongating RNA polymerase define exons and introns

Differential elongation rates of RNA polymerase II (RNAP) have been posited to be a critical determinant for pre-mRNA splicing. Molecular dissection of mechanisms coupling transcription elongation rate with splicing requires knowledge of instantaneous RNAP elongation velocity at exon and introns. However, only average RNAP elongation rates over large genomic distances can be inferred with current approaches, and local instantaneous velocities of the elongating RNA polymerase across endogenous genomic regions remain difficult to determine at sufficient resolution to enable detailed kinetic analysis of RNAP at exons. In order to overcome these challenges and to investigate kinetic features of RNAP elongation at genomic scale, we have employed global nuclear run-on sequencing (GRO-seq) method to infer changes in local RNAP elongation rates across the human genome, as changes in the residence time of RNAP. Using this approach, we have investigated functional coupling between the changes in local pattern of RNAP elongation rate at the exons and their general expression level, as inferred by sequencing of mRNAs (mRNA-seq). Our genomic level analyses reveal acceleration of RNAP at lowly expressed exons and confirm the previously reported deceleration of RNAP at highly expressed exons, suggesting variable local velocities of elongating RNAP that are potentially associated with different inclusion or exclusion rates of exons across the human genome.\n\nAUTHOR SUMMARYUnderstanding the mechanisms that enable high precision recognition and splicing of exons is fundamental to many aspects of human development and disease. Emerging data suggest that the speed of the elongating RNA polymerase affects pre-mRNA splicing; however, systematic genomic investigation of RNAP elongation speed and pre-mRNA have been lacking. Using a recently developed method for detecting synthesized nascent RNAs, we have inferred variable elongation rates of RNA polymerase II (RNAP) that are associated with included exons, introns and excluded exons, across the human genome. From this analysis, we have identified acceleration of RNAP at exons as a major determinant of exon exclusion across the genome, while confirming previous studies showing deceleration of RNAP at included exons.

Genomics

AC-PCA: simultaneous dimension reduction and adjustment for confounding variation

Dimension reduction methods are commonly applied to high-throughput biological datasets. However, the results can be hindered by confounding factors, either biologically or technically originated. In this study, we extend Principal Component Analysis to propose AC-PCA for simultaneous dimension reduction and adjustment for confounding variation. We show that AC-PCA can adjust for a) variations across individual donors present in a human brain exon array dataset, and b) variations of different species in a model organism ENCODE RNA-Seq dataset. Our approach is able to recover the anatomical structure of neocortical regions, and to capture the shared variation among species during embryonic development. For gene selection purposes, we extend AC-PCA with sparsity constraints, and propose and implement an efficient algorithm. The methods developed in this paper can also be applied to more general settings.

Bioinformatics

Integrative tissue-specific functional annotations in the human genome provide novel insights on many complex traits and improve signal prioritization in genome wide association studies

Extensive efforts have been made to understand genomic function through both experimental and computational approaches, yet proper annotation still remains challenging, especially in non-coding regions. In this manuscript, we introduce GenoSkyline, an unsupervised learning framework to predict tissue-specific functional regions through integrating high-throughput epigenetic annotations. GenoSkyline successfully identified a variety of non-coding regulatory machinery including enhancers, regulatory miRNA, and hypomethylated transposable elements in extensive case studies. Integrative analysis of GenoSkyline annotations and results from genome-wide association studies (GWAS) led to novel biological insights on the etiologies of a number of human complex traits. We also explored using tissue-specific functional annotations to prioritize GWAS signals and predict relevant tissue types for each risk locus. Brain and blood-specific annotations led to better prioritization performance for schizophrenia than standard GWAS p-values and non-tissue-specific annotations. As for coronary artery disease, heart-specific functional regions was highly enriched of GWAS signals, but previously identified risk loci were found to be most functional in other tissues, suggesting a substantial proportion of still undetected heart-related loci. In summary, GenoSkyline annotations can guide genetic studies at multiple resolutions and provide valuable insights in understanding complex diseases. GenoSkyline is available at http://genocanyon.med.yale.edu/GenoSkyline.

Bioinformatics

GenoWAP: Post-GWAS Prioritization Through Integrated Analysis of Genomic Functional Annotation

MotivationGenome-wide association study (GWAS) has been a great success in the past decade. However, significant challenges still remain in both identifying new risk loci and interpreting results. Bonferroni-corrected significance level is known to be conservative, leading to insufficient statistical power when the effect size is moderate at risk locus. Complex structure of linkage disequilibrium also makes it challenging to separate causal variants from nonfunctional ones in large haplotype blocks.\n\nResultsWe describe GenoWAP, a post-GWAS prioritization method that integrates genomic functional annotation and GWAS test statistics. The effectiveness of GenoWAP is demonstrated through its applications to Crohns disease and schizophrenia using the largest studies available, where highly ranked loci show substantially stronger signals in the whole dataset after prioritization based on a subset of samples. At the single nucleotide polymorphism (SNP) level, top ranked SNPs after prioritization have both higher replication rates and consistently stronger enrichment of eQTLs. Within each risk locus, GenoWAP is also able to distinguish functional sites from groups of correlated SNPs.\n\nAvailability and ImplementationGenoWAP is freely available on the web at http://genocanyon.med.yale.edu/GenoWAP

Bioinformatics

A Statistical Framework to Predict Functional Non-Coding Regions in the Human Genome Through Integrated Analysis of Annotation Data

Identifying functional regions in the human genome is a major goal in human genetics. Great efforts have been made to functionally annotate the human genome either through computational predictions, such as genomic conservation, or high-throughput experiments, such as the ENCODE project. These efforts have resulted in a rich collection of functional annotation data of diverse types that need to be jointly analyzed for integrated interpretation and annotation. Here we present GenoCanyon, a whole-genome annotation method that performs unsupervised statistical learning using 22 computational and experimental annotations thereby inferring the functional potential of each position in the human genome. With GenoCanyon, we are able to predict many of the known functional regions. The ability of predicting functional regions as well as its generalizable statistical framework makes GenoCanyon a unique and powerful tool for whole-genome annotation. The GenoCanyon web server is available at http://genocanyon.med.yale.edu

Bioinformatics

Mediated pleiotropy between psychiatric disorders and autoimmune disorders revealed by integrative analysis of multiple GWAS

Epidemiological observations and molecular-level experiments have indicated that brain disorders in the realm of psychiatry may be influenced by immune dysregulation. However, the degree of genetic overlap between immune disorders and psychiatric disorders has not been well established. We investigated this issue by integrative analysis of genome-wide association studies (GWAS) of 18 complex human traits/diseases (five psychiatric disorders, seven autoimmune disorders, and others) and multiple genomewide annotation resources (Central nervous system genes, immune-related expressionquantitative trait loci (eQTL) and DNase I hypertensive sites from 98 cell-lines). We detected pleiotropy in 24 of the 35 psychiatric-autoimmune disorder pairs, with statistical significance as strong as p=3.9e-285 (schizophrenia-rheumatoid arthritis). Strong enrichment (>1.4 fold) of immune-related eQTL was observed in four psychiatric disorders. Genomic regions responsible for pleiotropy between psychiatric disorders and autoimmune disorders were detected. The MHC region on chromosome 6 appears to be the most important (and it was indeed previously noted (1-3) as a confluence between schizophrenia and immune disorder risk regions), with many other regions, such as cytoband 1p13.2. We also found that most alleles shared between schizophrenia and Crohns disease have the same effect direction, with similar trend found for other disorder pairs, such as bipolar-Crohns disease. Our results offer a novel birds-eye view of the genetic relationship and demonstrate strong evidence for mediated pleiotropy between psychiatric disorders and autoimmune disorders. Our findings might open new routes for prevention and treatment strategies for these disorders based on a new appreciation of the importance of immunological mechanisms in mediating risk.

Genomics