Search bioRxivSearch

Biology subjects

Pasaniuc, B.

Publications and source records attributed to Pasaniuc, B..

10 recordsLinked to original sources

Large-scale transcriptome-wide association study identifies new prostate cancer risk regions

Although genome-wide association studies (GWAS) for prostate cancer (PrCa) have identified more than 100 risk regions, most of the risk genes at these regions remain largely unknown. Here, we integrate the largest PrCa GWAS (N=142,392) with gene expression measured in 45 tissues (N=4,458), including normal and tumor prostate, to perform a multi-tissue transcriptomewide association study (TWAS) for PrCa. We identify 235 genes at 87 independent 1Mb regions associated with PrCa risk, 9 of which are regions with no genome-wide significant SNP within 2Mb. 24 genes are significant in TWAS only for alternative splicing models in prostate tumor thus supporting the hypothesis of splicing driving risk for continued oncogenesis. Finally, we use a Bayesian probabilistic approach to estimate credible sets of genes containing the causal gene at pre-defined level; this reduced the list of 235 associations to 120 genes in the 90% credible set. Overall, our findings highlight the power of integrating expression with PrCa GWAS to identify novel risk loci and prioritize putative causal genes at known risk loci.

genetics

Multi-Tissue Transcriptome-Wide Association Studies Identify 21 Novel Candidate Susceptibility Genes for High Grade Serous Epithelial Ovarian Cancer

Genome-wide association studies (GWASs) have identified about 30 different susceptibility loci associated with high grade serous ovarian cancer (HGSOC) risk. We sought to identify potential susceptibility genes by integrating the risk variants at these regions with genetic variants impacting gene expression and splicing of nearby genes. We compiled gene expression and genotyping data from 2,169 samples for 6 different HGSOC-relevant tissue types. We integrated these data with GWAS data from 13,037 HGSOC cases and 40,941 controls, and performed a transcriptome-wide association study (TWAS) across >70,000 significantly heritable gene/exon features. We identified 24 transcriptome-wide significant associations for 14 unique genes, plus 90 significant exon-level associations in 20 unique genes. We implicated multiple novel genes at risk loci, e.g. LRRC46 at 19q21.32 (TWAS P=1x10-9) and a PRC1 splicing event (TWAS P=9x10-8) which was splice-variant specific and exhibited no eQTL signal. Functional analyses in HGSOC cell lines found evidence of essentiality for GOSR2, INTS1, KANSL1 and PRC1; with the latter gene showing levels of essentiality comparable to that of MYC. Overall, gene expression and splicing events explained 41% of SNP-heritability for HGSOC (s.e. 11%, P=2.5x10-4), implicated at least one target gene for 6/13 distinct genome-wide significant regions and revealed 2 known and 26 novel candidate susceptibility genes for HGSOC.\n\nSTATEMENT OF SIGNIFICANCEFor many ovarian cancer risk regions, the target genes regulated by germline genetic variants are unknown. Using expression data from >2,100 individuals, this study identified novel associations of genes and splicing variants with ovarian cancer risk; with transcriptional variation now explaining over one-third of the SNP-heritability for this disease.

genomics

Phenotype-specific enrichment of Mendelian disorder genes near GWAS regions across 62 complex traits

Although recent studies provide evidence for a common genetic basis between complex traits and Mendelian disorders, a thorough quantification of their overlap in a phenotype-specific manner remains elusive. Here, we quantify the overlap of genes identified through large-scale genome-wide association studies (GWAS) for 62 complex traits and diseases with genes known to cause 20 broad categories of Mendelian disorders. We identify a significant enrichment of phenotypically-matched Mendelian disorder genes in GWAS gene sets. Further, we observe elevated GWAS effect sizes near phenotypically-matched Mendelian disorder genes. Finally, we report examples of GWAS variants localized at the transcription start site or physically interacting with the promoters of phenotypically-matched Mendelian disorder genes. Our results are consistent with the hypothesis that genes that are disrupted in Mendelian disorders are dysregulated by noncoding variants in complex traits, and demonstrate how leveraging findings from related Mendelian disorders and functional genomic datasets can prioritize genes that are putatively dysregulated by local and distal non-coding GWAS variants.

genetics

Haplotype-based eQTL mapping finds evidence for complex gene regulatory regions poorly tagged by marginal SNPs

MotivationExpression quantitative trait loci (eQTLs), variations in the genome that impact gene expression, are identified through eQTL studies that test for a relationship between single nucleotide polymorphisms (SNPs) and gene expression levels. These studies typically assume an underlying additive model. Non-additive tests have been proposed, but are limited due to the increase in the multiple testing burden and are potentially biased by filtering criteria that relies on marginal association data. Here we propose using combinations of short haplotypes instead of SNPs as predictors for gene expression. Essentially, this method looks for genomic regions where haplotypes have different effect sizes. The differences in effect can be due to multiple genetic architectures such as a single SNP, a burden of rare SNPs, multiple SNPs with independent effect or multiple SNPs with an interaction effect occurring on the same haplotype.\n\nResultsSimulations show that when haplotypes, rather than SNPs, are assigned non-zero effect sizes, our method has increased power compared to the marginal SNP method. In the GEUVADIS gene expression data, our method finds 101 more eGenes than the marginal method (5,202 vs. 5,101). The methods do not have full overlap in the eGenes that they find. Of the 5,202 eGenes found by our method, 707 are not found by the marginal method--even though it has a lower significance threshold. This indicates that many genes have regulatory architectures that are not well tagged by marginal SNPs and demonstrates the need to better model alternative archi-tectures.

genetics

A unifying framework for joint trait analysis under a non-infinitesimal model

MotivationA large proportion of risk regions identified by genome-wide association studies (GWAS) are shared across multiple diseases and traits. Understanding whether this clustering is due to sharing of causal variants or chance colocalization can provide insights into shared etiology of complex traits and diseases.\n\nResultsIn this work, we propose a flexible, unifying framework to quantify the overlap between a pair of traits called UNITY (Unifying Non-Infinitesimal Trait analYsis). We formulate a Bayesian generative model that relates the overlap between pairs of traits to GWAS summary statistic data under a non-infinitesimal genetic architecture underlying each trait. We propose a Metropolis-Hastings sampler to compute the posterior density of the genetic overlap parameters in this model. We validate our method through comprehensive simulations and analyze summary statistics from height and BMI GWAS to show that it produces estimates consistent with the known genetic makeup of both traits.\n\nAvailabilityThe UNITY software is made freely available to the research community at: https://github.com/bogdanlab/UNITY\n\nContactruthjohnson@ucla.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Probabilistic fine-mapping of transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) using predicted expression have identified thousands of genes whose locally-regulated expression is associated to complex traits and diseases. In this work, we show that linkage disequilibrium (LD) among SNPs induce significant gene-trait associations at non-causal genes as a function of the overlap between eQTL weights used in expression prediction. We introduce a probabilistic framework that models the induced correlation among TWAS signals to assign a probability for every gene in the risk region to explain the observed association signal while controlling for pleiotropic SNP effects and unmeasured causal expression. Importantly, our approach remains accurate when expression data for causal genes are not available in the causal tissue by leveraging expression prediction from other tissues. Our approach yields credible-sets of genes containing the causal gene at a nominal confidence level (e.g., 90%) that can be used to prioritize and select genes for functional assays. We illustrate our approach using an integrative analysis of lipids traits where our approach prioritizes genes with strong evidence for causality.

genetics

Assessing the genetic effect mediated through gene expression from summary eQTL and GWAS data

Integrating genome-wide association (GWAS) and expression quantitative trait locus (eQTL) data into transcriptome-wide association studies (TWAS) based on predicted expression can boost power to detect novel disease loci or pinpoint the susceptibility gene at a known disease locus. However, it is often the case that multiple eQTL genes colocalize at disease loci, making the identification of the true susceptibility gene challenging, due to confounding through linkage disequilibrium (LD). To distinguish between true susceptibility genes (where the genetic effect on phenotype is mediated through expression) and colocalization due to LD, we examine an extension of the Mendelian Randomization Egger regression method that allows for LD while only requiring summary association data for both GWAS and eQTL. We derive the standard TWAS approach in the context of Mendelian Randomization and show in simulations that the standard TWAS does not control Type I error for causal gene identification when eQTLs have pleiotropic or LD-confounded effects on disease. In contrast, LD Aware MR-Egger regression can control Type I error in this case while attaining similar power as other methods in situations where these provide valid tests. However, when the direct effects of genetic variants on traits are correlated with the eQTL associations, all of the methods we examined including LD Aware MR-Egger regression can have inflated Type I error. We illustrate these methods by integrating gene expression within a recent large-scale breast cancer GWAS to provide guidance on susceptibility gene identification.

epidemiology

Leveraging polygenic functional enrichment to improve GWAS power

Functional genomics data has the potential to increase GWAS power by identifying SNPs that have a higher prior probability of association. Here, we introduce a method that leverages polygenic functional enrichment to incorporate coding, conserved, regulatory and LD-related genomic annotations into association analyses. We show via simulations with real genotypes that the method, Functionally Informed Novel Discovery Of Risk loci (FINDOR), correctly controls the false-positive rate at null loci and attains a 9-38% increase in the number of independent associations detected at causal loci, depending on trait polygenicity and sample size. We applied FINDOR to 27 independent complex traits and diseases from the interim UK Biobank release (average N=130K). Averaged across traits, we attained a 13% increase in genome-wide significant loci detected (including a 20% increase for disease traits) compared to un-weighted raw p-values that do not use functional data. We replicated the novel loci in independent UK Biobank and non-UK Biobank data, yielding a highly statistically significant replication slope (0.66-0.69) in each case. Finally, we applied FINDOR to the full UK Biobank release (average N=416K), attaining smaller relative improvements (consistent with simulations) but larger absolute improvements, detecting an additional 583 GWAS loci. In conclusion, leveraging functional enrichment using our method robustly increases GWAS power.

genetics

A Bayesian Framework for Multiple Trait Colocalization from Summary Association Statistics

MotivationMost genetic variants implicated in complex diseases by genome-wide association studies (GWAS) are non-coding, making it challenging to understand the causative genes involved in disease. Integrating external information such as quantitative trait locus (QTL) mapping of molecular traits (e.g., expression, methylation) is a powerful approach to identify the subset of GWAS signals explained by regulatory effects. In particular, expression QTLs (eQTLs) help pinpoint the responsible gene among the GWAS regions that harbor many genes, while methylation QTLs (mQTLs) help identify the epigenetic mechanisms that impact gene expression which in turn affect disease risk. In this work we propose multiple-trait-coloc (moloc), a Bayesian statistical framework that integrates GWAS summary data with multiple molecular QTL data to identify regulatory effects at GWAS risk loci.\n\nResultsWe applied moloc to schizophrenia (SCZ) and eQTL/mQTL data derived from human brain tissue and identified 52 candidate genes that influence SCZ through methylation. Our method can be applied to any GWAS and relevant functional data to help prioritize disease associated genes.\n\nAvailabilitymoloc is available for download as an R package (https://github.com/clagiamba/moloc). We also developed a web site to visualize the biological findings (icahn.mssm.edu/moloc). The browser allows searches by gene, methylation probe, and scenario of interest.\n\nContactclaudia.giambartolomei@gmail.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

genomics

Local genetic correlation gives insights into the shared genetic architecture of complex traits

Although genetic correlations between complex traits provide valuable insights into epidemiological and etiological studies, a precise quantification of which genomic regions contribute to the genome-wide genetic correlation is currently lacking. Here, we introduce{rho} -HESS, a technique to quantify the correlation between pairs of traits due to genetic variation at a small region in the genome. Our approach only requires GWAS summary data and makes no distributional assumption on the causal variant effects sizes while accounting for linkage disequilibrium (LD) and overlapping GWAS samples. We analyzed large-scale GWAS summary data across 35 complex traits, and identified 27 genomic regions that contribute significantly to the genetic correlation among these traits. Notably, we find 7 genomic regions that contribute to the genetic correlation of 12 pairs of traits that show negligible genome-wide correlation, further showcasing the power of local genetic correlation analyses. Finally, we leverage the distribution of local genetic correlations across the genome to assign putative direction of causality for 15 pairs of traits.

genetics