Search bioRxivSearch

Biology subjects

Strauch, K.

Publications and source records attributed to Strauch, K..

4 recordsLinked to original sources

Evaluation of the causal effect of fibrinogen on incident coronary heart disease via Mendelian randomization

BackgroundFibrinogen is an essential hemostatic factor and cardiovascular disease risk factor. Early attempts at evaluating the causal effect of fibrinogen on coronary heart disease (CHD) and myocardial infraction (MI) using Mendelian randomization (MR) used single variant approaches, and did not take advantage of recent genome-wide association studies (GWAS) or multi-variant, pleiotropy robust MR methodologies.\n\nMethods and FindingsWe evaluated evidence for a causal effect of fibrinogen on both CHD and MI using MR. We used both an allele score approach and pleiotropy robust MR models. The allele score was composed of 38 fibrinogen-associated variants from recent GWAS. Initial analyses using the allele score incorporated data from 11 European-ancestry prospective cohorts to examine incidence CHD and MI. We also applied 2 sample MR methods with data from a prevalent CHD and MI GWAS. Results are given in terms of the hazard ratio (HR) or odds ratio (OR), depending on the study design, and associated 95% confidence interval (CI).\n\nIn single variant analyses no causal effect of fibrinogen on CHD or MI was observed. In multi-variant analyses using incidence CHD cases and the allele score approach, the estimated causal effect (HR) of a 1 g/L higher fibrinogen concentration was 1.62 (CI = 1.12, 2.36) when using incident cases and the allele score approach. In 2 sample MR analyses that accounted for pleiotropy, the causal estimate (OR) was reduced to 1.18 (CI = 0.98, 1.42) and 1.09 (CI = 0.89, 1.33) in the 2 most precise (smallest CI) models, out of 4 models evaluated. In the 2 sample MR analyses for MI, there was only very weak evidence of a causal effect in only 1 out of 4 models.\n\nConclusionsA small causal effect of fibrinogen on CHD is observed using multi-variant MR approaches which account for pleiotropy, but not single variant MR approaches. Taken together, results indicate that even with large sample sizes and multi-variant approaches MR analyses still cannot exclude the null when estimating the causal effect of fibrinogen on CHD, but that any potential causal effect is likely to be much smaller than observed in epidemiological studies.\n\nAuthor SummaryInitial Mendelian Randomization (MR) analyses of the causal effect of fibrinogen on coronary heart disease (CHD) utilized single variants and did not take advantage of modern, multivariant approaches. This manuscript provides an important update to these initial analyses by incorporating larger sample sizes and employing multiple, modern multi-variant MR approaches to account for pleiotropy. We used incident cases to perform a MR study of the causal effect of fibrinogen on incident CHD and the nested outcome of myocardial infarction (MI) using an allele score approach. Then using data from a case-control genome-wide association study for CHD and MI we performed two sample MR analyses with multiple, pleiotropy robust approaches. Overall, the results indicated that associations between fibrinogen and CHD in observational studies are likely upwardly biased from any underlying causal effect. Single variant MR approaches show little evidence of a causal effect of fibrinogen on CHD or MI. Multi-variant MR analyses of fibrinogen on CHD indicate there may be a small positive effect, however this result needs to be interpreted carefully as the 95% confidence intervals were still consistent with a null effect. Multi-variant MR approaches did not suggest evidence of even a small causal effect of fibrinogen on MI.

genetics

Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies

BACKGROUNDUntargeted mass spectrometry (MS)-based metabolomics data often contain missing values that reduce statistical power and can introduce bias in epidemiological studies. However, a systematic assessment of the various sources of missing values and strategies to handle these data has received little attention. Missing data can occur systematically, e.g. from run day-dependent effects due to limits of detection (LOD); or it can be random as, for instance, a consequence of sample preparation.\n\nMETHODSWe investigated patterns of missing data in an MS-based metabolomics experiment of serum samples from the German KORA F4 cohort (n = 1750). We then evaluated 31 imputation methods in a simulation framework and biologically validated the results by applying all imputation approaches to real metabolomics data. We examined the ability of each method to reconstruct biochemical pathways from data-driven correlation networks, and the ability of the method to increase statistical power while preserving the strength of established genetically metabolic quantitative trait loci.\n\nRESULTSRun day-dependent LOD-based missing data accounts for most missing values in the metabolomics dataset. Although multiple imputation by chained equations (MICE) performed well in many scenarios, it is computationally and statistically challenging. K-nearest neighbors (KNN) imputation on observations with variable pre-selection showed robust performance across all evaluation schemes and is computationally more tractable.\n\nCONCLUSIONMissing data in untargeted MS-based metabolomics data occur for various reasons. Based on our results, we recommend that KNN-based imputation is performed on observations with variable pre-selection since it showed robust results in all evaluation schemes.\n\nKey messagesO_LIUntargeted MS-based metabolomics data show missing values due to both batch-specific LOD-based and non-LOD-based effects.\nC_LIO_LIStatistical evaluation of multiple imputation methods was conducted on both simulated and real datasets.\nC_LIO_LIBiological evaluation on real data assessed the ability of imputation methods to preserve statistical inference of biochemical pathways and correctly estimate effects of genetic variants on metabolite levels.\nC_LIO_LIKNN-based imputation on observations with variable pre-selection and K = 10 showed robust performance for all data scenarios across all evaluation schemes.\nC_LI

systems biology

Network based conditional genome wide association analysis of human metabolomics

BackgroundGenome-wide association studies (GWAS) have identified hundreds of loci influencing complex human traits, however, their biological mechanism of action remains mostly unknown. Recent accumulation of functional genomics ( omics) including metabolomics data opens up opportunities to provide a new insight into the functional role of specific changes in the genome. Functional genomic data are characterized by high dimensionality, presence of (strong) statistical dependencies between traits, and, potentially, complex genetic control. Therefore, analysis of such data asks for development of specific statistical genetic methods.\n\nResultsWe propose a network-based, conditional approach to evaluate the impact of genetic variants on omics phenotypes (conditional GWAS, cGWAS). For each trait of interest, based on biological network, we select a set of other traits to be used as covariates in GWAS. The network could be reconstructed either from biological pathway databases or directly from the data. We evaluated our approach using data from a population-based KORA study (n=1,784, 1.7 M SNPs) with measured metabolomics data (151 metabolites) and demonstrated that our approach allows for identification of up to five additional loci not detected by conventional GWAS. We show that this gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.\n\nConclusionsWe demonstrate that in context of metabolomics network-based, conditional genome-wide association analysis is able to dramatically increase power of identification of loci with specific contra-intuitive pleiotropic architecture. Our method has modest computational costs, can utilize summary level GWAS data, and is applicable to other omics data types. We anticipate that application of our method to new and existing data sets will facilitate progress in understanding genetic bases of control of molecular and complex phenotypes.\n\nShort abstractWe propose a network-based, conditional approach for genome-wide analysis of multivariate omics phenotypes. Our methods can incorporate prior biological knowledge about biological pathways from external sources. We evaluated our approach using metabolomics data and demonstrated that our approach has bigger power and allows for identification of additional loci. We show that gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.

genomics

Connecting genetic risk to disease endpoints through the human blood plasma proteome

Genome-wide association studies (GWAS) with intermediate phenotypes, like changes in metabolite and protein levels, provide functional evidence for mapping disease associations and translating them into clinical applications. However, although hundreds of genetic risk variants have been associated with complex disorders, the underlying molecular pathways often remain elusive. Associations with intermediate traits across multiple chromosome locations are key in establishing functional links between GWAS-identified risk-variants and disease endpoints. Here, we describe a GWAS performed with a highly multiplexed aptamer-based affinity proteomics platform. We quantified associations between protein level changes and gene variants in a German cohort and replicated this GWAS in an Arab/Asian cohort. We identified many independent, SNP-protein associations, which represent novel, inter-chromosomal links, related to autoimmune disorders, Alzheimer's disease, cardiovascular disease, cancer, and many other disease endpoints. We integrated this information into a genome-proteome network, and created an interactive web-tool for interrogations. Our results provide a basis for new approaches to pharmaceutical and diagnostic applications.

genetics