Search bioRxivSearch

Biology subjects

Gieger, C.

Publications and source records attributed to Gieger, C..

8 recordsLinked to original sources

Alcohol consumption is associated with widespread changes in blood DNA methylation: analysis of cross-sectional and longitudinal data

BackgroundDNA methylation may be one of the mechanisms by which alcohol consumption is associated with the risk of disease. We conducted a large-scale, cross-sectional, genome-wide DNA methylation association study of alcohol consumption and a longitudinal analysis of repeated measurements taken several years apart.\n\nMethodsUsing the Illumina Infinium HumanMethylation450 BeadChip, DNA methylation measures were determined using baseline peripheral blood samples from 5,606 adult Melbourne Collaborative Cohort Study (MCCS) participants. For a subset of 1,088 of them, these measures were repeated using blood samples collected at follow-up, a median of 11 years later. Associations between alcohol intake and blood DNA methylation were assessed using linear mixed-effects regression models adjusted for batch effects and potential confounders. Independent data from the LOLIPOP (N=4,042) and KORA (N=1,662) cohorts were used to replicate associations discovered in the MCCS.\n\nResultsCross-sectional analyses identified 1,414 CpGs associated with alcohol intake at P<10-7, 1,243 of which had not been reported previously. Of these 1,243 novel associations, 1,078 were replicated (P<0.05) using LOLIPOP and KORA data. Using the MCCS data, we also replicated (P<0.05) 403 of 518 associations that had been reported previously. Interaction analyses suggested that associations were stronger for women, non-smokers, and participants genetically predisposed to consume less alcohol. Of the 1,414 CpGs, 530 were differentially methylated (P<0.05) in former compared with current drinkers. Longitudinal associations between the change in alcohol intake and the change in methylation were observed for 513 of the 1,414 cross-sectional associations.\n\nConclusionOur study indicates that, for middle-aged and older adults, alcohol intake is associated with widespread changes in DNA methylation across the genome. Longitudinal analyses showed that the methylation status of alcohol-associated CpGs may change with changes in alcohol consumption.

epidemiology

Defining the genetic control of human blood plasma N-glycome using genome-wide association study

Glycosylation is a common post-translational modification of proteins. It is known, that glycans are directly involved in the pathophysiology of every major disease. Defining genetic factors altering glycosylation may provide a basis for novel approaches to diagnostic and pharmaceutical applications. Here, we report a genome-wide association study of the human blood plasma N-glycome composition in up to 3811 people. We discovered and replicated twelve loci. This allowed us to demonstrate a clear overlap in genetic control between total plasma and IgG glycosylation. Majority of loci contained genes that encode enzymes directly involved in glycosylation (FUT3/FUT6, FUT8, B3GAT1, ST6GAL1, B4GALT1, ST3GAL4, MGAT3, and MGAT5). We, however, also found loci that are likely to reflect other, more complex, aspects of plasma glycosylation process. Functional genomic annotation suggested the role of DERL3, which potentially highlights the role of glycoprotein degradation pathway, and such transcription factor as IKZF1.

genetics

Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies

BACKGROUNDUntargeted mass spectrometry (MS)-based metabolomics data often contain missing values that reduce statistical power and can introduce bias in epidemiological studies. However, a systematic assessment of the various sources of missing values and strategies to handle these data has received little attention. Missing data can occur systematically, e.g. from run day-dependent effects due to limits of detection (LOD); or it can be random as, for instance, a consequence of sample preparation.\n\nMETHODSWe investigated patterns of missing data in an MS-based metabolomics experiment of serum samples from the German KORA F4 cohort (n = 1750). We then evaluated 31 imputation methods in a simulation framework and biologically validated the results by applying all imputation approaches to real metabolomics data. We examined the ability of each method to reconstruct biochemical pathways from data-driven correlation networks, and the ability of the method to increase statistical power while preserving the strength of established genetically metabolic quantitative trait loci.\n\nRESULTSRun day-dependent LOD-based missing data accounts for most missing values in the metabolomics dataset. Although multiple imputation by chained equations (MICE) performed well in many scenarios, it is computationally and statistically challenging. K-nearest neighbors (KNN) imputation on observations with variable pre-selection showed robust performance across all evaluation schemes and is computationally more tractable.\n\nCONCLUSIONMissing data in untargeted MS-based metabolomics data occur for various reasons. Based on our results, we recommend that KNN-based imputation is performed on observations with variable pre-selection since it showed robust results in all evaluation schemes.\n\nKey messagesO_LIUntargeted MS-based metabolomics data show missing values due to both batch-specific LOD-based and non-LOD-based effects.\nC_LIO_LIStatistical evaluation of multiple imputation methods was conducted on both simulated and real datasets.\nC_LIO_LIBiological evaluation on real data assessed the ability of imputation methods to preserve statistical inference of biochemical pathways and correctly estimate effects of genetic variants on metabolite levels.\nC_LIO_LIKNN-based imputation on observations with variable pre-selection and K = 10 showed robust performance for all data scenarios across all evaluation schemes.\nC_LI

systems biology

Fine-mapping of an expanded set of type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps

We aggregated genome-wide genotyping data from 32 European-descent GWAS (74,124 T2D cases, 824,006 controls) imputed to high-density reference panels of >30,000 sequenced haplotypes. Analysis of {small tilde}27M variants ({small tilde}21M with minor allele frequency [MAF]<5%), identified 243 genome-wide significant loci (p<5x10-8; MAF 0.02%-50%; odds ratio [OR] 1.04-8.05), 135 not previously-implicated in T2D-predisposition. Conditional analyses revealed 160 additional distinct association signals (p<10-5) within the identified loci. The combined set of 403 T2D-risk signals includes 56 low-frequency (0.5%[&le;]MAF<5%) and 24 rare (MAF<0.5%) index SNPs at 60 loci, including 14 with estimated allelic OR>2. Forty-one of the signals displayed effect-size heterogeneity between BMI-unadjusted and adjusted analyses. Increased sample size and improved imputation led to substantially more precise localisation of causal variants than previously attained: at 51 signals, the lead variant after fine-mapping accounted for >80% posterior probability of association (PPA) and at 18 of these, PPA exceeded 99%. Integration with islet regulatory annotations enriched for T2D association further reduced median credible set size (from 42 variants to 32) and extended the number of index variants with PPA>80% to 73. Although most signals mapped to regulatory sequence, we identified 18 genes as human validated therapeutic targets through coding variants that are causal for disease. Genome wide chip heritability accounted for 18% of T2D-risk, and individuals in the 2.5% extremes of a polygenic risk score generated from the GWAS data differed >9-fold in risk. Our observations highlight how increases in sample size and variant diversity deliver enhanced discovery and single-variant resolution of causal T2D-risk alleles, and the consequent impact on mechanistic insights and clinical translation.

genomics

An exact and efficient score test for variance components models

Testing for the existence of variance components in linear mixed models is a fundamental task in many applicative fields. In statistical genetics, the score test has recently become instrumental in the task of testing an association between a set of genetic markers and a phenotype. With few markers, this amounts to set-based variance component tests, which attempt to increase power in association studies by aggregating weak individual effects. When the entire genome is considered, it allows testing for the heritability of a phenotype, defined as the proportion of phenotypic variance explained by genetics. In the popular score-based Sequence Kernel Association Test (SKAT) method, the assumed distribution of the score test statistic is uncalibrated in small samples, with a correction being computationally expensive. This may cause severe inflation or deflation of p-values, even when the null hypothesis is true. Here, we characterize the conditions under which this discrepancy holds, and show it may occur also in large real datasets, such as a dataset from the Wellcome Trust Case Control Consortium 2 (n=13,950) study, and in particular when the individuals in the sample are unrelated. In these cases the SKAT approximation tends to be highly over-conservative and therefore underpowered. To address this limitation, we suggest an efficient method to calculate exact p-values for the score test in the case of a single variance component and a continuous response vector, which can speed up the analysis by orders of magnitude. Our results enable fast and accurate application of the score test in heritability and in set-based association tests. Our method is available in http://github.com/cozygene/RL-SKAT.

genetics

Genome-wide Association Study Of Plasma Proteins Identifies Putatively Causal Genes, Proteins, And Pathways For Cardiovascular Disease

Identifying genetic variants associated with circulating protein concentrations (pQTLs) and integrating them with variants from genome-wide association studies (GWAS) may illuminate the proteomes causal role in disease and bridge a GWAS knowledge gap for hitherto unexplained SNP-disease associations. We conducted GWAS of 71 high-value proteins for cardiovascular disease in 6,861 Framingham Heart Study participants followed by external replication. We comprehensively mapped thousands of pQTLs, including functional annotations and clinical-trait associations, and created an integrated plasma-protein-QTL searchable database. We next identified 15 proteins with pQTLs coinciding with coronary heart disease (CHD)-related variants from GWAS or tested causal for CHD by Mendelian randomization; most of these proteins were associated with new-onset cardiovascular disease events in Framingham participants with long-term follow-up. Identifying pQTLs and integrating them with GWAS results yields insights into genes, proteins, and pathways that may be causally associated with disease and can serve as therapeutic targets for treatment and prevention.

epidemiology

Network based conditional genome wide association analysis of human metabolomics

BackgroundGenome-wide association studies (GWAS) have identified hundreds of loci influencing complex human traits, however, their biological mechanism of action remains mostly unknown. Recent accumulation of functional genomics ( omics) including metabolomics data opens up opportunities to provide a new insight into the functional role of specific changes in the genome. Functional genomic data are characterized by high dimensionality, presence of (strong) statistical dependencies between traits, and, potentially, complex genetic control. Therefore, analysis of such data asks for development of specific statistical genetic methods.\n\nResultsWe propose a network-based, conditional approach to evaluate the impact of genetic variants on omics phenotypes (conditional GWAS, cGWAS). For each trait of interest, based on biological network, we select a set of other traits to be used as covariates in GWAS. The network could be reconstructed either from biological pathway databases or directly from the data. We evaluated our approach using data from a population-based KORA study (n=1,784, 1.7 M SNPs) with measured metabolomics data (151 metabolites) and demonstrated that our approach allows for identification of up to five additional loci not detected by conventional GWAS. We show that this gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.\n\nConclusionsWe demonstrate that in context of metabolomics network-based, conditional genome-wide association analysis is able to dramatically increase power of identification of loci with specific contra-intuitive pleiotropic architecture. Our method has modest computational costs, can utilize summary level GWAS data, and is applicable to other omics data types. We anticipate that application of our method to new and existing data sets will facilitate progress in understanding genetic bases of control of molecular and complex phenotypes.\n\nShort abstractWe propose a network-based, conditional approach for genome-wide analysis of multivariate omics phenotypes. Our methods can incorporate prior biological knowledge about biological pathways from external sources. We evaluated our approach using metabolomics data and demonstrated that our approach has bigger power and allows for identification of additional loci. We show that gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.

genomics

Connecting genetic risk to disease endpoints through the human blood plasma proteome

Genome-wide association studies (GWAS) with intermediate phenotypes, like changes in metabolite and protein levels, provide functional evidence for mapping disease associations and translating them into clinical applications. However, although hundreds of genetic risk variants have been associated with complex disorders, the underlying molecular pathways often remain elusive. Associations with intermediate traits across multiple chromosome locations are key in establishing functional links between GWAS-identified risk-variants and disease endpoints. Here, we describe a GWAS performed with a highly multiplexed aptamer-based affinity proteomics platform. We quantified associations between protein level changes and gene variants in a German cohort and replicated this GWAS in an Arab/Asian cohort. We identified many independent, SNP-protein associations, which represent novel, inter-chromosomal links, related to autoimmune disorders, Alzheimer's disease, cardiovascular disease, cancer, and many other disease endpoints. We integrated this information into a genome-proteome network, and created an interactive web-tool for interrogations. Our results provide a basis for new approaches to pharmaceutical and diagnostic applications.

genetics