Search bioRxivSearch

Biology subjects

Hemani, G.

Publications and source records attributed to Hemani, G..

28 records · Page 2Linked to original sources

Automating Mendelian randomization through machine learning to construct a putative causal map of the human phenome

A major application for genome-wide association studies (GWAS) has been the emerging field of causal inference using Mendelian randomization (MR), where the causal effect between a pair of traits can be estimated using only summary level data. MR depends on SNPs exhibiting vertical pleiotropy, where the SNP influences an outcome phenotype only through an exposure phenotype. Issues arise when this assumption is violated due to SNPs exhibiting horizontal pleiotropy. We demonstrate that across a range of pleiotropy models, instrument selection will be increasingly liable to selecting invalid instruments as GWAS sample sizes continue to grow. Methods have been developed in an attempt to protect MR from different patterns of horizontal pleiotropy, and here we have designed a mixture-of-experts machine learning framework (MR-MoE 1.0) that predicts the most appropriate model to use for any specific causal analysis, improving on both power and false discovery rates. Using the approach, we systematically estimated the causal effects amongst 2407 phenotypes. Almost 90% of causal estimates indicated some level of horizontal pleiotropy. The causal estimates are organised into a publicly available graph database (http://eve.mrbase.org), and we use it here to highlight the numerous challenges that remain in automated causal inference.

epidemiology

Investigating the role of insulin in increased adiposity: Bi-directional Mendelian randomization study

Insulin may serve as a key causal agent which regulates fat accumulation in the body. Here we assessed the causal relationship between fasting insulin and adiposity using publicly-available results from two large-scale genome-wide association studies for body mass index and fasting insulin levels in a two-sample, bidirectional Mendelian Randomized approach. This approach is only valid on the condition that the two instruments are independent of one another. In analysis excluding overlapping loci, there was an increase of 0.20 (0.17, 0.23) log pmol/L fasting insulin per SD increase in BMI (P= 2.80 x 10-36), while there was a null effect of fasting insulin on BMI, with a 0.01 (-0.39, 0.38) SD decrease in BMI per log pmol/L increase in fasting insulin (P= 0.98). Furthermore, a high degree of heterogeneity in the causal estimates was obtained from the insulin-related variants, which may be attributed to varying mechanisms of action of the insulin-associated variants. Results were largely consistent when an Egger regression technique and weighted median and mode estimators were applied. Findings suggest that the positive correlation between adiposity and fasting insulin levels are at least in part explained by the causal effect of adiposity on increasing insulin, rather than vice versa.

epidemiology

PhenoSpD: an atlas of phenotypic correlations and a multiple testing correction for the human phenome

BackgroundIdentifying phenotypic correlations between complex traits and diseases can provide useful etiological insights. Restricted access to individual-level phenotype data makes it difficult to estimate large-scale phenotypic correlation across the human phenome. State-of-the-art methods, metaCCA and LD score regression, provide an alternative approach to estimate phenotypic correlation using genome-wide association study (GWAS) summary statistics.\n\nResultsHere, we present an integrated R toolkit, PhenoSpD, to 1) apply metaCCA (or LD score regression) to estimate phenotypic correlations using GWAS summary statistics; and 2) to utilize the estimated phenotypic correlations to inform correction of multiple testing for complex human traits using the spectral decomposition of matrices (SpD). The simulations suggest it is possible to estimate phenotypic correlation using samples with only a partial overlap, but as overlap decreases correlations will attenuate towards zero and multiple testing correction will be more stringent than in perfectly overlapping samples. In a case study, PhenoSpD using GWAS results suggested 324.4 independent tests among 452 metabolites, which is close to the 296 independent tests estimated using true phenotypic correlation. We further applied PhenoSpD to estimated 7,503 pair-wise phenotypic correlations among 123 metabolites using GWAS summary statistics from Kettunen et al. and PhenoSpD suggested 44.9 number of independent tests for theses metabolites.\n\nConclusionPhenoSpD integrates existing methods and provides a simple and conservative way to reduce dimensionality for complex human traits using GWAS summary statistics, which is particularly valuable for post-GWAS analysis of complex molecular traits.\n\nAvailabilityR code and documentation for PhenoSpD V1.0.0 is available online (https://github.com/MRCIEU/PhenoSpD).

bioinformatics

Partitioning Phenotypic Variance Due To Parent-Of-Origin Effects Using Genomic Relatedness Matrices

Introduction Introduction Methods Statistical Methods Results Discussion References Parent-of-origin effects (POEs) describe the phenomenon in which the effects of alleles depend upon their parental origin. POEs imply that heterozygote individuals have phenotypes which are distributed differently depending upon which of their alleles were maternally and paternally transmitted (Guilmatre and Sharp 2012; Lawson et al. 2013). The extreme case of POEs is polar overdominance, where the two heterozygotes' phenotypes differ in distribution but the two homozygotes share the same distribution (Hoggart et al. 2014). Imprinting, a phenomenon in which one parent's allele is not expressed, is probably the most widely studied example of POE (Peters 2 ...

genetics

Causal epigenome-wide association study identifies CpG sites that influence cardiovascular disease risk

The extent to which genetic influences on complex traits and disease are mediated by changes in DNA methylation levels has not been systematically explored. We developed an analytical framework that integrates genetic fine mapping and Mendelian randomization with epigenome-wide association studies to evaluate the causal relationships between methylation levels and 14 cardiovascular disease traits.\n\nWe identified 10 genetic loci known to influence proximal DNA methylation which were also associated with cardiovascular traits (P < 3.83x10-08). Bivariate fine mapping suggested that the individual variants responsible for the observed effects on cardiovascular traits at the ABO, ADCY3, ADIPOQ, APOA1 and IL6R loci were likely mediated through changes in DNA methylation. Causal effect estimates on cardiovascular traits ranged between 0.109-0.992 per standard deviation change in DNA methylation and were replicated using results from large-scale consortia.\n\nFunctional informatics suggests that the causal variants and CpG sites identified in this study were enriched for histone mark peaks in adipose tissue and gene promoter regions. Integrating our results with expression quantitative trait loci data we provide evidence that variation at these regulatory regions is likely to also influence gene expression at these loci.

epidemiology

Meffil: efficient normalisation and analysis of very large DNA methylation samples

AbstractO_ST_ABSBackgroundC_ST_ABSTechnological advances in high throughput DNA methylation microarrays have allowed dramatic growth of a new branch of epigenetic epidemiology. DNA methylation datasets are growing ever larger in terms of the number of samples profiled, the extent of genome coverage, and the number of studies being meta-analysed. Novel computational solutions are required to efficiently handle these data.\n\nMethodsWe have developed meffil, an R package designed to quality control, normalize and perform epigenome-wide association studies (EWAS) efficiently on large samples of Illumina Infinium HumanMethylation450 and MethylationEPIC BeadChip microarrays. We tested meffil by applying it to 6000 450k microarrays generated from blood collected for two different datasets, Accessible Resource for Integrative Epigenomic Studies (ARIES) and The Genetics of Overweight Young Adults (GOYA) study.\n\nResultsA complete reimplementation of functional normalization minimizes computational memory requirements to 5% of that required by other R packages, without increasing running time. Incorporating fixed and random effects alongside functional normalization, and automated estimation of functional normalisation parameters reduces technical variation in DNA methylation levels, thus reducing false positive associations and improving power. We also demonstrate that the ability to normalize datasets distributed across physically different locations without sharing any biologically-based individual-level data may reduce heterogeneity in meta-analyses of epigenome-wide association studies. However, we show that when batch is perfectly confounded with cases and controls functional normalization is unable to prevent spurious associations.\n\nConclusionsmeffil is available online (https://github.com/perishky/meffil/) along with tutorials covering typical use cases.

bioinformatics

The causal effect of educational attainment on Alzheimer’s disease: A two-sample Mendelian randomization study

BackgroundObservational evidence suggests that higher educational attainment is protective for Alzheimers disease (AD). It is unclear whether this association is causal or confounded by demographic and socioeconomic characteristics. We examined the causal effect of educational attainment on AD in a two-sample MR framework.\n\nMethodsWe extracted all available effect estimates of the 74 single nucleotide polymorphisms (SNPs) associated with years of schooling from the largest genome-wide association study (GWAS) of educational attainment (N=293,723) and the GWAS of AD conducted by the International Genomics of Alzheimers Project (n=17,008 AD cases and 37,154 controls). SNP-exposure and SNP-outcome coefficients were combined using an inverse variance weighted approach, providing an estimate of the causal effect of each SD increase in years of schooling on AD. We also performed appropriate sensitivity analyses examining the robustness of causal effect estimates to the various assumptions and conducted simulation analyses to examine potential survival bias of MR analyses.\n\nFindingsWith each SD increase in years of schooling (3.51 years), the odds of AD were, on average, reduced by approximately one third (odds ratio= 0.63, 95% confidence interval [CI]: 0.48 to 0.83, p<0.001). Causal effect estimates were consistent when using causal methods with varying MR assumptions or different sets of SNPs for educational attainment, lending confidence to the magnitude and direction of effect in our main findings. There was also no evidence of survival bias in our study.\n\nInterpretationOur findings support a causal role of educational attainment on AD, whereby an additional [~]3.5 years of schooling reduces the odds of AD by approximately one third.

epidemiology

The role of glycaemic and lipid risk factors in mediating the effect of BMI on coronary heart disease: A two-step, two-sample Mendelian randomization study

BackgroundThe extent to which effects of BMI on coronary heart disease (CHD) are mediated by gylcaemic and lipid risk factors is unclear.\n\nMethodsWe used two-sample Mendelian randomization to determine the causal effect of: (i) BMI on CHD (60,801 cases; 123, 504 controls), type 2 diabetes (T2DM; 34,840 cases; 114,981 controls), fasting glucose (n=46,186), insulin (n=38,238), HbA1c (n=46,368), LDL-cholesterol (LDL-C), HDL-cholesterol (HDL-C) and triglycerides (n=188,577); (ii) glycaemic and lipids traits on CHD; and (iii) extent to which these traits mediated any effect of BMI on CHD.\n\nFindingsOne standard deviation (SD) increase in BMI (~ 4.5kg/m2) increased CHD (odds ratio=1.45 (95% confidence interval (CI): 1.27, 1.66)) and T2DM (1.96 (1.35, 2.83)), and levels of fasting glucose (0.07mmol/l (95%CI 0.03, 0.11)), HbA1c (0.05% (95%CI 0.01, 0.08)), fasting insulin (0.18log pmol/l (95%CI 0.14, 0.22)) and triglycerides (0.20 SD (95%CI 0.14, 0.26)), and lowered levels of HDL-C (-0.23 SD (95%CI -0.32, -0.15)). BMI was not causally related to LDL-C. After accounting for potential pleiotropy, triglycerides, HbA1c and T2DM were causally related to CHD. The BMI-CHD effect reduced from 1.45 to 1.16 (95%CI 0.99, 1.36) and to 1.36 (95%CI 1.19, 1.57) with genetic adjustment for triglycerides or HbA1c respectively, and to 1.09 (95%CI 0.94, 1.27) with adjustment for both.\n\nInterpretationIncreased triglyceride levels and poor glycaemic control appear to mediate much of the effect of BMI on CHD.\n\nFundingEuropean Research Council (669545), European Union (733206), China Medical Board (CMB_2015/16), Conselho Nacional de Desenvolvimento Cientifico e Tecnologico and UK Medical Research Council (MC_UU_12013/5).

epidemiology

Orienting The Causal Relationship Between Imprecisely Measured Traits Using Genetic Instruments

Inference of the causal structure that induces correlations between two traits can be achieved by combining genetic associations with a mediation-based approach, as is done in the causal inference test (CIT) and others. However, we show that measurement error in the phenotypes can lead to mediation-based approaches inferring the wrong causal direction, and that increasing sample sizes has the adverse effect of increasing confidence in the wrong answer. Here we introduce an extension to Mendelian randomisation, a method that uses genetic associations in an instrumentation framework, that enables inference of the causal direction between traits, with some advantages. First, it is less susceptible to bias in the presence of measurement error; second, it is more statistically efficient; third, it can be performed using only summary level data from genome-wide association studies; and fourth, its sensitivity to measurement error can be evaluated. We apply the method to infer the causal direction between DNA methylation and gene expression levels. Our results demonstrate that, in general, DNA methylation is more likely to be the causal factor, but this result is highly susceptible to bias induced by systematic differences in measurement error between the platforms. We emphasise that, where possible, implementing MR and appropriate sensitivity analyses alongside other approaches such as CIT is important to triangulate reliable conclusions about causality.

systems biology

MR-Base: a platform for systematic causal inference across the phenome using billions of genetic associations

Published genetic associations can be used to infer causal relationships between phenotypes, bypassing the need for individual-level genotype or phenotype data. We have curated complete summary data from 1094 genome-wide association studies (GWAS) on diseases and other complex traits into a centralised database, and developed an analytical platform that uses these data to perform Mendelian randomization (MR) tests and sensitivity analyses (MR-Base, http://www.mrbase.org). Combined with curated data of published GWAS hits for phenomic measures, the MR-Base platform enables millions of potential causal relationships to be evaluated. We use the platform to predict the impact of lipid lowering on human health. While our analysis provides evidence that reducing LDL-cholesterol, lipoprotein(a) or triglyceride levels reduce coronary disease risk, it also suggests causal effects on a number of other non-vascular outcomes, indicating potential for adverse-effects or drug repositioning of lipid-lowering therapies.

epidemiology