Search bioRxivSearch

Biology subjects

Burgess, S.

Publications and source records attributed to Burgess, S..

9 recordsLinked to original sources

Selecting causal risk factors from high-throughput experiments using multivariable Mendelian randomization

Modern high-throughput experiments provide a rich resource to investigate causal determinants of disease risk. Mendelian randomization (MR) is the use of genetic variants as instrumental variables to infer the causal effect of a specific risk factor on an outcome. Multivariable MR is an extension of the standard MR framework to consider multiple potential risk factors in a single model. However, current implementations of multivariable MR use standard linear regression and hence perform poorly with many risk factors.\n\nHere, we propose a novel approach to two-sample multivariable MR based on Bayesian model averaging (MR-BMA) that scales to high-throughput experiments. In a realistic simulation study, we show that MR-BMA can detect true causal risk factors even when the candidate risk factors are highly correlated. We illustrate MR-BMA by analysing publicly-available summarized data on metabolites to prioritise likely causal biomarkers for age-related macular degeneration.

genetics

Disentangling genetic overlap between Attention-Deficit/Hyperactivity Disorder, literacy and language

Interpreting polygenic overlap between ADHD and both literacy- and language-related impairments is challenging as genetic confounding can bias associations. Here, we investigate evidence for links between polygenic ADHD risk and multiple literacy- and language-related abilities (LRAs), assessed in UK children (N[&le;]5,919), conditional on genetic effects shared with educational attainment (EA). Genome-wide summary statistics on clinical ADHD and years-of-schooling were obtained from large consortia (N[&le;]326,041). ADHD-polygenic scores (ADHD-PGS) were inversely associated with LRAs in ALSPAC, most consistently with reading-related abilities, and explained [&le;]1.6% phenotypic variation. Polygenic links were then dissected into both genetic effects shared with and independent of EA using multivariable regressions (MVR), analogous to Mendelian Randomization approaches accounting for mediating effects. Conditional on EA, polygenic ADHD risk remained associated with multiple literacy-related skills, phonemic awareness and verbal intelligence, but not language-related skills such as listening comprehension and non-word repetition. Pooled reading performance showed the strongest overlap with ADHD independent of EA. Using conservative ADHD-instruments (P-threshold<5x10-8) this corresponded to a 0.35 decrease in Z-scores per log-odds in ADHD-liability (P=9.2x10-5). Using subthreshold ADHD-instruments (P-threshold<0.0015), these associations had lower magnitude, but higher predictive accuracy, with a 0.03 decrease in Z-scores (P=1.4x10-6). Polygenic ADHD-effects shared with EA were of equal strength and at least equal magnitude compared to those independent of EA, for all LRAs studied, and only detectable using subthreshold instruments. Thus, ADHD-related polygenic links are highly susceptible to genetic confounding, concealing an ADHD-specific association profile that primarily involves reading-related impairments, but few language-related problems.

genetics

Birthweight, Type 2 Diabetes and Cardiovascular Disease: Addressing the Barker Hypothesis with Mendelian randomization

BackgroundLow birthweight (BW) has been associated with a higher risk of hypertension, type 2 diabetes (T2D) and cardiovascular disease (CVD) in epidemiological studies. The Barker hypothesis posits that intrauterine growth restriction resulting in lower BW is causal for these diseases, but causality and mechanisms are difficult to infer from observational studies. Mendelian randomization (MR) is a new tool to address this important question.\n\nMethodsWe performed regression analyses to assess associations of self-reported BW with CVD and T2D in 237,631 individuals from the UK Biobank, a large population-based cohort study aged 40-69 years recruited across UK in 2006-2010. Further, we assessed the causal relationship of such associations using the two- sample MR approach, estimating the causal effect by contrasting the SNP effects on the exposure with the SNP effects on the outcome using independent publicly available genome-wide association datasets.\n\nResultsIn the observational analyses, BW showed strong inverse associations with systolic and diastolic blood pressure ({beta}, -0.83 and -0.26; per raw unit in outcomes and SD change in BW; 95% CI, -0.90, -0.75 and -0.31, -0.22, respectively), T2D (odds ratio [OR], 0.83; 95% CI, 0.79, 0.87), lipid-lowering treatment (OR, 0.84; 95% CI, 0.81, 0.86) and CAD (hazard ratio [HR] 0.85; 95% CI, 0.78, 0.94); while the associations with adult body mass index (BMI) and body fat ({beta}, 0.04 and 0.02; per SD change in outcomes and BW; 95% CI, 0.03, 0.04 and 0.01, 0.02, respectively) were positive. The MR analyses indicated inverse causal associations of BW with low density lipoprotein cholesterol, 2-hour glucose, CAD and T2D, and positive causal association with BMI; but no associations with blood pressure. Sensitivity analyses and robust MR methods provided consistent results and indicated no horizontal pleiotropy.\n\nConclusionOur study indicates that lower BW is causally and directly related with increased susceptibility to CAD and T2D in adulthood. This causal relationship is not mediated by adult obesity or hypertension.

genetics

Improving on a modal-based estimation method: model averaging for consistent and efficient estimation in Mendelian randomization when a plurality of candidate instruments are valid

BackgroundA robust method for Mendelian randomization does not require all genetic variants to be valid instruments to give consistent estimates of a causal parameter. Several such methods have been developed, including a mode-based estimation method giving consistent estimates if a plurality of genetic variants are valid instruments; that is, there is no larger subset of invalid instruments estimating the same causal parameter than the subset of valid instruments.\n\nMethodsWe here develop a model averaging method that gives consistent estimates under the same plurality of valid instruments assumption. The method considers a mixture distribution of estimates derived from each subset of genetic variants. The estimates are weighted such that subsets with more genetic variants receive more weight, unless variants in the subset have heterogeneous causal estimates, in which case that subset is severely downweighted. The mode of this mixture distribution is the causal estimate. This heterogeneity-penalized model averaging method has several technical advantages over the previously proposed mode-based estimation method.\n\nResultsThe heterogeneity-penalized model averaging method outperformed the mode-based estimation in terms of effciency and outperformed other robust methods in terms of Type 1 error rate in an extensive simulation analysis. The proposed method suggests two distinct mechanisms by which inflammation affects coronary heart disease risk, with subsets of variants suggesting both positive and negative causal effects.\n\nConclusionsThe heterogeneity-penalized model averaging method is an additional robust method for Mendelian randomization with excellent theoretical and practical properties, and can reveal features in the data such as the presence of multiple causal mechanisms. (249 words)\n\nKey messagesO_LIWe propose a heterogeneity-penalized model averaging method that gives consistent causal estimates if a weighted plurality of the genetic variants are valid instruments.\nC_LIO_LIThe method calculates causal estimates based on all subsets of genetic variants, and upweights subsets containing several genetic variants with similar causal estimates.\nC_LIO_LIThe method is asymptotically effcient and does not rely on bootstrapping to obtain a confidence interval, nor is the confidence interval constrained to be symmetric.\nC_LIO_LIIn particular, the confidence interval can include multiple disjoint intervals, suggesting the presence of multiple causal mechanisms by which the risk factor influences the outcome.\nC_LIO_LIThe method can incorporate biological knowledge to upweight the contribution of genetic variants with stronger plausibility of being valid instruments.\nC_LI

genetics

Dissecting causal pathways using Mendelian randomization with summarized genetic data: application to age at menarche and risk of breast cancer

Mendelian randomization is the use of genetic variants as instrumental variables to estimate causal effects of risk factors on outcomes. The total causal effect of a risk factor is the change in the outcome resulting from intervening on the risk factor. This total causal effect may potentially encompass multiple mediating mechanisms. For a proposed mediator, the direct effect of the risk factor is the change in the outcome resulting from a change in the risk factor keeping the mediator constant. A difference between the total effect and the direct effect indicates that the causal pathway from the risk factor to the outcome acts at least in part via the mediator (an indirect effect). Here, we show that Mendelian randomization estimates of total and direct effects can be obtained using summarized data on genetic associations with the risk factor, mediator, and outcome, potentially from different data sources. We perform simulations to test the validity of this approach when there is unmeasured confounding and/or bidirectional effects between the risk factor and mediator. We illustrate this method using the relationship between age at menarche and risk of breast cancer, with body mass index (BMI) as a potential mediator. We show an inverse direct causal effect of age at menarche on risk of breast cancer (independent of BMI) and a positive indirect effect via BMI. In conclusion, multivariable Mendelian randomization using summarized genetic data provides a rapid and accessible analytic strategy that can be undertaken using publicly-available data to better understand causal mechanisms.

epidemiology

Consequences Of Natural Perturbations In The Human Plasma Proteome

Proteins are the primary functional units of biology and the direct targets of most drugs, yet there is limited knowledge of the genetic factors determining inter-individual variation in protein levels. Here we reveal the genetic architecture of the human plasma proteome, testing 10.6 million DNA variants against levels of 2,994 proteins in 3,301 individuals. We identify 1,927 genetic associations with 1,478 proteins, a 4-fold increase on existing knowledge, including trans associations for 1,104 proteins. To understand consequences of perturbations in plasma protein levels, we introduce an approach that links naturally occurring genetic variation with biological, disease, and drug databases. We provide insights into pathogenesis by uncovering the molecular effects of disease-associated variants. We identify causal roles for protein biomarkers in disease through Mendelian randomization analysis. Our results reveal new drug targets, opportunities for matching existing drugs with new disease indications, and potential safety concerns for drugs under development.

genomics

Semiparametric methods for estimation of a non-linear exposure-outcome relationship using instrumental variables with application to Mendelian randomization

Mendelian randomization, the use of genetic variants as instrumental variables (IV), can test for and estimate the causal effect of an exposure on an outcome. Most IV methods assume that the function relating the exposure to the expected value of the outcome (the exposure-outcome relationship) is linear. However, in practice this assumption may not hold. Indeed, often the primary question of interest is to assess the shape of this relationship. We present two novel IV methods for investigating the shape of the exposure-outcome relationship: a fractional polynomial method and a piecewise linear method. We divide the population into strata using the exposure distribution, and estimate a causal effect, referred to as a localized average causal effect (LACE), in each stratum of population. The fractional polynomial method performs meta-regression on these LACE estimates. The piecewise linear method estimates a continuous piecewise linear function, the gradient of which is the LACE estimate in each stratum. Both methods were demonstrated in a simulation study to estimate the true exposure-outcome relationship well, particularly when the relationship was a fractional polynomial (for the fractional polynomial method) or was piecewise linear (for the piecewise linear method). The methods were used to investigate the shape of relationship of body mass index with systolic blood pressure and diastolic blood pressure.\n\nAvailability and implementation: https://github.com/jrs95/nlmr

genetics

MR-Base: a platform for systematic causal inference across the phenome using billions of genetic associations

Published genetic associations can be used to infer causal relationships between phenotypes, bypassing the need for individual-level genotype or phenotype data. We have curated complete summary data from 1094 genome-wide association studies (GWAS) on diseases and other complex traits into a centralised database, and developed an analytical platform that uses these data to perform Mendelian randomization (MR) tests and sensitivity analyses (MR-Base, http://www.mrbase.org). Combined with curated data of published GWAS hits for phenomic measures, the MR-Base platform enables millions of potential causal relationships to be evaluated. We use the platform to predict the impact of lipid lowering on human health. While our analysis provides evidence that reducing LDL-cholesterol, lipoprotein(a) or triglyceride levels reduce coronary disease risk, it also suggests causal effects on a number of other non-vascular outcomes, indicating potential for adverse-effects or drug repositioning of lipid-lowering therapies.

epidemiology

Power calculator for instrumental variable analysis in pharmacoepidemiology

BackgroundInstrumental variable analysis, for example with physicians prescribing preferences as an instrument for medications issued in primary care, is an increasingly popular method in the field of pharmacoepidemiology. Existing power calculators for studies using instrumental variable analysis, such as Mendelian randomisation power calculators, do not allow for the structure of research questions in this field. This is because the analysis in pharmacoepidemiology will typically have stronger instruments and detect larger causal effects than in other fields. Consequently, there is a need for dedicated power calculators for pharmacoepidemiological research.\n\nMethods and resultsThe formula for calculating the power of a study using instrumental variable analysis in the context of pharmacoepidemiology is derived before being validated by a simulation study. The formula is applicable for studies using a single binary instrument to analyse the causal effect of a binary exposure on a continuous outcome. A web application is provided for the implementation of the formula by others.\n\nConclusionsThe statistical power of instrumental variable analysis in pharmacoepidemiological studies to detect a clinically meaningful treatment effect is an important consideration. Research questions in this field have distinct structures that must be accounted for when calculating power.\n\nFUNDING STATEMENTThis work was supported by the Perros Trust and the Integrative Epidemiology Unit. The Integrative Epidemiology Unit is supported by the Medical Research Council and the University of Bristol [grant number MC_UU_12013/9]. Stephen Burgess is supported by a post-doctoral fellowship from the Wellcome Trust [100114].\n\nKey MessagesO_LIResearch questions using instrumental variable analysis in pharmacoepidemiology have distinct structures that have previously not been catered for by instrumental variable analysis power calculators.\nC_LIO_LIPower can be calculated for studies using a single binary instrument to analyse the causal effect of a binary exposure on a continuous outcome in the context of pharmacoepidemiology using the presented formula and online power calculator.\nC_LIO_LIThe use of this power calculator will allow investigators to determine whether a pharmacoepidemiology study is likely to detect clinically meaningful treatment effects prior to the studys commencement.\nC_LI

epidemiology