Search bioRxivSearch

Biology subjects

Manolis Kellis

Publications and source records attributed to Manolis Kellis.

9 recordsLinked to original sources

Distant regulatory effects of genetic variation in multiple human tissues

Understanding the genetics of gene regulation provides information on the cellular mechanisms through which genetic variation influences complex traits. Expression quantitative trait loci, or eQTLs, are enriched for polymorphisms that have been found to be associated with disease risk. While most analyses of human data has focused on regulation of expression by nearby variants (cis-eQTLs), distal or trans-eQTLs may have broader effects on the transcriptome and important phenotypic consequences, necessitating a comprehensive study of the effects of genetic variants on distal gene transcription levels. In this work, we identify trans-eQTLs in the Genotype Tissue Expression (GTEx) project data1, consisting of 449 individuals with RNA-sequencing data across 44 tissue types. We find 81 genes with a trans-eQTL in at least one tissue, and we demonstrate that trans-eQTLs are more likely than cis-eQTLs to have effects specific to a single tissue. We evaluate the genomic and functional properties of trans-eQTL variants, identifying strong enrichment in enhancer elements and Piwi-interacting RNA clusters. Finally, we describe three tissue-specific regulatory loci underlying relevant disease associations: 9q22 in thyroid that has a role in thyroid cancer, 5q31 in skeletal muscle, and a previously reported master regulator near KLF14 in adipose. These analyses provide a comprehensive characterization of trans-eQTLs across human tissues, which contribute to an improved understanding of the tissue-specific cellular mechanisms of regulatory genetic variation.

Genomics

Local genetic effects on gene expression across 44 human tissues

Expression quantitative trait locus (eQTL) mapping provides a powerful means to identify functional variants influencing gene expression and disease pathogenesis. We report the identification of cis-eQTLs from 7,051 post-mortem samples representing 44 tissues and 449 individuals as part of the Genotype-Tissue Expression (GTEx) project. We find a cis-eQTL for 88% of all annotated protein-coding genes, with one-third having multiple independent effects. We identify numerous tissue-specific cis-eQTLs, highlighting the unique functional impact of regulatory variation in diverse tissues. By integrating large-scale functional genomics data and state-of-the-art fine-mapping algorithms, we identify multiple features predictive of tissue-specific and shared regulatory effects. We improve estimates of cis-eQTL sharing and effect sizes using allele specific expression across tissues. Finally, we demonstrate the utility of this large compendium of cis-eQTLs for understanding the tissue-specific etiology of complex traits, including coronary artery disease. The GTEx project provides an exceptional resource that has improved our understanding of gene regulation across tissues and the role of regulatory variation in human genetic diseases.

Genomics

RiVIERA-beta: Joint Bayesian inference of risk variants and tissue-specific epigenomic enrichments across multiple complex human diseases

Genome wide association studies (GWAS) provide a powerful approach for uncovering disease-associated variants in human, but fine-mapping the causal variants remains a challenge. This is partly remedied by prioritization of disease-associated variants that overlap GWAS-enriched epigenomic annotations. Here, we introduce a new Bayesian model RiVIERA-beta (Risk Variant Inference using Epigenomic Reference Annotations) for inference of driver variants by modelling summary statistics p-values in Beta density function across multiple traits using hundreds of epigenomic annotations. In simulation, RiVIERA-beta promising power in detecting causal variants and causal annotations, the multi-trait joint inference further improved the detection power. We applied RiVIERA-beta to model the existing GWAS summary statistics of 9 autoimmune diseases and Schizophrenia by jointly harnessing the potential causal enrichments among 848 tissue-specific epigenomics annotations from ENCODE/Roadmap consortium covering 127 cell/tissue types and 8 major epigenomic marks. RiVIERA-beta identified meaningful tissue-specific enrichments for enhancer regions defined by H3K4me1 and H3K27ac for Blood T-Cell specifically in the 9 autoimmune diseases and Brain-specific enhancer activities exclusively in Schizophrenia. Moreover, the variants from the 95% credible sets exhibited high conservation and enrichments for GTEx whole-blood eQTLs located within transcription-factor-binding-sites and DNA-hypersensitive-sites. Furthermore, joint modeling the nine immune traits by simultaneously inferring and exploiting the underlying epigenomic correlation between traits further improved the functional enrichments compared to single-trait models.

Bioinformatics

RiVIERA-MT: A Bayesian model to infer risk variants in related traits using summary statistics and functional genomic annotations

Dissecting the physiological circuitry underlying diverse human complex traits associated with heritable common mutations is an ongoing effort. The primary challenge involves identifying the relevant cell types and the causal variants among the vast majority of the associated mutations in the noncoding regions. To address this challenge, we developed an efficient probabilistic framework. First, we propose a sparse group-guided learning algorithm to infer cell-type-specific enrichments. Second, we propose a fine-mapping Bayesian model that incorporates as Bayesian priors the sparse enrichments to infer risk variants. Using the proposed framework to analyze 32 complex human traits revealed meaningful tissue-specific epigenomic enrichments indicative of the relevant disease pathologies. The prioritized variants exhibit prominent tissue-specific epigenomic signatures and significant enrichments for eQTL and conserved elements. Together, we demonstrate the general benefits of the proposed integrative framework in elucidating meaningful tissue-specific epigenomic elements from large-scale correlated annotations and the implicated functional variants for future experimental interrogation.

Bioinformatics

Evolutionary dynamics of abundant stop codon readthrough in Anopheles and Drosophila

Translational stop codon readthrough was virtually unknown in eukaryotic genomes until recent developments in comparative genomics and new experimental techniques revealed evidence of readthrough in hundreds of fly genes and several human, worm, and yeast genes. Here, we use the genomes of 21 species of Anopheles mosquitoes and improved comparative techniques to identify evolutionary signatures of conserved, functional readthrough of 353 stop codons in the malaria vector, Anopheles gambiae, and 51 additional Drosophila melanogaster stop codons, with several cases of double and triple readthrough including readthrough of two adjacent stop codons, supporting our earlier prediction of abundant readthrough in pancrustacea genomes. Comparisons between Anopheles and Drosophila allow us to transcend the static picture provided by single-clade analysis to explore the evolutionary dynamics of abundant readthrough. We find that most differences between the readthrough repertoires of the two species are due to readthrough gain or loss in existing genes, rather than to birth of new genes or to gene death; that RNA structures are sometimes gained or lost while readthrough persists; and that readthrough is more likely to be lost at TAA and TAG stop codons. We also determine which characteristic properties of readthrough predate readthrough and which are clade-specific. We estimate that there are more than 600 functional readthrough stop codons in A. gambiae and 900 in D. melanogaster. We find evidence that readthrough is used to regulate peroxisomal targeting in two genes. Finally, we use the sequenced centipede genome to refine the phylogenetic extent of abundant readthrough.

Molecular Biology

Evidence of a recombination rate valley in human regulatory domains

Human recombination rate varies greatly, but the forces shaping it remain incompletely understood. Here, we study the relationship between recombination rate and gene-regulatory domains defined by a gene and its linked control elements. We define these links using methylation quantitative trait loci (meQTLs), expression quantitative trait loci (eQTLs), chromatin conformation, and correlated activity across cell types. Each link type shows a \"recombination valley\" of significantly-reduced recombination rate compared to control regions, indicating preferential co-inheritance of genes and linked regulatory elements as a single unit. This recombination valley is most pronounced for gene-regulatory domains of early embryonic developmental genes, housekeeping genes, and constitutive regulatory elements, which are known to show increased evolutionary constraint across species. Recombination valleys show increased DNA methylation, reduced double-stranded break initiation, and increased repair efficiency, specifically in the lineage leading to the germ line, providing a potential molecular mechanism facilitating their maintenance by exclusion of recombination events.

Genomics

Functional enrichments of disease variants across thousands of independent loci in eight diseases

For most complex traits, known genetic associations only explain a small fraction of the narrow sense heritability prompting intense debate on the genetic basis of complex traits. Joint analysis of all common variants together explains much of this missing heritability and reveals that large numbers of weakly associated loci are enriched in regulatory regions, but fails to identify specific regions or biological pathways. Here, we use epigenomic annotations across 127 tissues and cell types to investigate weak regulatory associations, the specific enhancers they reside in, their downstream target genes, their upstream regulators, and the biological pathways they disrupt in eight common diseases. We show weak associations are significantly enriched in disease-relevant regulatory regions across thousands of independent loci. We develop methods to control for LD between weak associations and overlap between annotations. We show that weak non-coding associations are additionally enriched in relevant biological pathways implicating additional downstream target genes and upstream disease-specific master regulators. Our results can help guide the discovery of biologically meaningful, but currently undetectable regulatory loci underlying a number of common diseases.

Genetics

Various localized epigenetic marks predict expression across 54 samples and reveal underlying chromatin state enrichments

Here, we predict gene expression from epigenetic features based on public data available through the Epigenome Roadmap Project [1]. This rich new dataset includes samples from primary tissues, which to our knowledge have not previously been studied in this context. Specifically, we used computational machine learning algorithms on five histone modifications to predict gene expression in a variety of samples. Our models reveal a high predictive accuracy, especially in cell cultures, with predictive ability dependent on sample type and anatomy. The relative importance of each histone mark feature varied across samples. We localized each histone mark signal to its relevant region, revealing that chromatin state enrichment varies greatly between histone marks. Our results provide several novel insights into epigenetic regulation of transcription in new contexts.

Genetics

Phylogenetic Identification and Functional Characterization of Orthologs and Paralogs across Human, Mouse, Fly, and Worm

Model organisms can serve the biological and medical community by enabling the study of conserved gene families and pathways in experimentally-tractable systems. Their use, however, hinges on the ability to reliably identify evolutionary orthologs and paralogs with high accuracy, which can be a great challenge at both small and large evolutionary distances. Here, we present a phylogenomics-based approach for the identification of orthologous and paralogous genes in human, mouse, fly, and worm, which forms the foundation of the comparative analyses of the modENCODE and mouse ENCODE projects. We study a median of 16,101 genes across 2 mammalian genomes (human, mouse), 12 Drosophila genomes, 5 Caenorhabditis genomes, and an outgroup yeast genome, and demonstrate that accurate inference of evolutionary relationships and events across these species must account for frequent gene-tree topology errors due to both incomplete lineage sorting and insufficient phylogenetic signal. Furthermore, we show that integration of two separate phylogenomic pipelines yields increased accuracy, suggesting that their sources of error are independent, and finally, we leverage the resulting annotation of homologous genes to study the functional impact of gene duplication and loss in the context of rich gene expression and functional genomic datasets of the modENCODE, mouse ENCODE, and human ENCODE projects.

Genomics