Search bioRxivSearch

Biology subjects

Suhre, K.

Publications and source records attributed to Suhre, K..

10 recordsLinked to original sources

Defining the genetic control of human blood plasma N-glycome using genome-wide association study

Glycosylation is a common post-translational modification of proteins. It is known, that glycans are directly involved in the pathophysiology of every major disease. Defining genetic factors altering glycosylation may provide a basis for novel approaches to diagnostic and pharmaceutical applications. Here, we report a genome-wide association study of the human blood plasma N-glycome composition in up to 3811 people. We discovered and replicated twelve loci. This allowed us to demonstrate a clear overlap in genetic control between total plasma and IgG glycosylation. Majority of loci contained genes that encode enzymes directly involved in glycosylation (FUT3/FUT6, FUT8, B3GAT1, ST6GAL1, B4GALT1, ST3GAL4, MGAT3, and MGAT5). We, however, also found loci that are likely to reflect other, more complex, aspects of plasma glycosylation process. Functional genomic annotation suggested the role of DERL3, which potentially highlights the role of glycoprotein degradation pathway, and such transcription factor as IKZF1.

genetics

MoDentify: a tool for phenotype-driven module identification in multilevel metabolomics networks

SummaryMetabolomics is an established tool to gain insights into (patho)physiological outcomes. Associations of metabolism with such outcomes are expected to span functional modules, which are defined as sets of correlating metabolites that are coordinately regulated. Moreover, these associations occur at different scales, from entire pathways to only a few metabolites, which is an aspect that has not been addressed by previous methods. Here we present MoDentify, a freely available R package to identify regulated modules in metabolomics networks at different layers of resolution. Importantly, MoDentify shows higher statistical power than classical association analysis. Moreover, the package offers direct visualization of results as interactive networks in Cytoscape. We present an application example using a complex, multifluid metabolomics dataset. Owing to its generic character, the method is widely applicable to any dataset with a phenotype variable, a data matrix, and optional pathway annotations.\n\nAvailability and ImplementationMoDentify is freely available from GitHub: https://github.com/krumsiek/MoDentify\n\nThe package vignette contains a detailed tutorial of the analysis workflow.\n\nContactjan.krumsiek@helmholtz-muenchen.de

systems biology

Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies

BACKGROUNDUntargeted mass spectrometry (MS)-based metabolomics data often contain missing values that reduce statistical power and can introduce bias in epidemiological studies. However, a systematic assessment of the various sources of missing values and strategies to handle these data has received little attention. Missing data can occur systematically, e.g. from run day-dependent effects due to limits of detection (LOD); or it can be random as, for instance, a consequence of sample preparation.\n\nMETHODSWe investigated patterns of missing data in an MS-based metabolomics experiment of serum samples from the German KORA F4 cohort (n = 1750). We then evaluated 31 imputation methods in a simulation framework and biologically validated the results by applying all imputation approaches to real metabolomics data. We examined the ability of each method to reconstruct biochemical pathways from data-driven correlation networks, and the ability of the method to increase statistical power while preserving the strength of established genetically metabolic quantitative trait loci.\n\nRESULTSRun day-dependent LOD-based missing data accounts for most missing values in the metabolomics dataset. Although multiple imputation by chained equations (MICE) performed well in many scenarios, it is computationally and statistically challenging. K-nearest neighbors (KNN) imputation on observations with variable pre-selection showed robust performance across all evaluation schemes and is computationally more tractable.\n\nCONCLUSIONMissing data in untargeted MS-based metabolomics data occur for various reasons. Based on our results, we recommend that KNN-based imputation is performed on observations with variable pre-selection since it showed robust results in all evaluation schemes.\n\nKey messagesO_LIUntargeted MS-based metabolomics data show missing values due to both batch-specific LOD-based and non-LOD-based effects.\nC_LIO_LIStatistical evaluation of multiple imputation methods was conducted on both simulated and real datasets.\nC_LIO_LIBiological evaluation on real data assessed the ability of imputation methods to preserve statistical inference of biochemical pathways and correctly estimate effects of genetic variants on metabolite levels.\nC_LIO_LIKNN-based imputation on observations with variable pre-selection and K = 10 showed robust performance for all data scenarios across all evaluation schemes.\nC_LI

systems biology

Accelerated lipid catabolism and autophagy are cancer survival mechanisms under inhibited glutaminolysis

Suppressing glutaminolysis does not always induce cancer cell death in glutamine-dependent tumors because cells may switch to alternative energy sources. To reveal compensatory metabolic pathways, we investigated the metabolome-wide cellular response to inhibited glutaminolysis. We conducted metabolic profiling in the triple-negative breast cancer cell line MB-MDA-231, treated with different dosages of glutaminase inhibitor C.968 at multiple time points. We found that multiple molecules involved in lipid catabolism responded directly to glutamate deficiency as a presumed compensation for energy deficit. Accelerated lipid catabolism, together with oxidative stress induced by glutaminolysis inhibition, triggered autophagy. We therefore simultaneously inhibited glutaminolysis and autophagy, which induced cancer cell death. Our study emphasizes the potential of non-targeted metabolomics to characterize and identify metabolic escape mechanisms contributing to cancer cell survival under treatment. Our findings add to the increasing evidence that combined inhibition of glutaminolysis and autophagy may be effective in glutamine-addicted cancers.

cancer biology

Genus-wide sequencing supports a two-locus model for sex-determination in Phoenix

The date palm tree is a commercially important member of the genus Phoenix whose 14 species are all dioecious with separate male and female individuals. Previous studies identified a multi-megabase region of the date palm genome linked to sex and showed that dioecy likely developed in Phoenix prior to speciation. To identify genes critical to sex determination we sequenced the genomes of 28 Phoenix trees representing all 14 species. Male-specific sequences were identified and extended using phased single molecule sequencing or BAC clones to distinguish X and Y alleles.\n\nHere we show that only four genes contain sequences conserved in all analyzed males, likely identifying the changes foundational to dioecy in Phoenix. The majority of these sequences show similarity to a single genomic locus in the closely related oil palm. CYP703 and GPAT3, two genes critical to male flower development in other monocots, appear fully deleted in females while maintained as single copy in males. A LOG-like gene appears translocated into the Y chromosome and a cytidine deaminase-like appears at the border of a chromosomal rearrangement. Our data supports a two-mutation model for the evolution from hermaphroditism to dioecy through a gynodioecious intermediate.

genomics

ProGeM: A framework for the prioritisation of candidate causal genes at molecular quantitative trait loci

Quantitative trait locus (QTL) mapping of molecular phenotypes such as metabolites, lipids, and proteins through genome-wide association studies (GWAS) represents a powerful means of highlighting molecular mechanisms relevant to human diseases. However, a major challenge of this approach is to identify the causal gene(s) at the observed QTLs. Here we present a framework for the \"Prioritisation of candidate causal Genes at Molecular QTLs\" (ProGeM), which incorporates biological domain-specific annotation data alongside genome annotation data from multiple repositories. We assessed the performance of ProGeM using a reference set of 227 previously reported and extensively curated metabolite QTLs. For 98% of these loci, the expert-curated gene was one of the candidate causal genes prioritised by ProGeM. Benchmarking analyses revealed that 69% of the causal candidates were nearest to the sentinel variant at the investigated molecular QTLs, indicating that genomic proximity is the most reliable indicator of \"true positive\" causal genes. In contrast, cis-gene expression QTL data led to three false positive candidate causal gene assignments for every one true positive assignment. We provide evidence that these conclusions also apply to other molecular phenotypes, suggesting that ProGeM is a powerful and versatile tool for annotating molecular QTLs. ProGeM is freely available via GitHub.

bioinformatics

Genome-wide Association Study Of Plasma Proteins Identifies Putatively Causal Genes, Proteins, And Pathways For Cardiovascular Disease

Identifying genetic variants associated with circulating protein concentrations (pQTLs) and integrating them with variants from genome-wide association studies (GWAS) may illuminate the proteomes causal role in disease and bridge a GWAS knowledge gap for hitherto unexplained SNP-disease associations. We conducted GWAS of 71 high-value proteins for cardiovascular disease in 6,861 Framingham Heart Study participants followed by external replication. We comprehensively mapped thousands of pQTLs, including functional annotations and clinical-trait associations, and created an integrated plasma-protein-QTL searchable database. We next identified 15 proteins with pQTLs coinciding with coronary heart disease (CHD)-related variants from GWAS or tested causal for CHD by Mendelian randomization; most of these proteins were associated with new-onset cardiovascular disease events in Framingham participants with long-term follow-up. Identifying pQTLs and integrating them with GWAS results yields insights into genes, proteins, and pathways that may be causally associated with disease and can serve as therapeutic targets for treatment and prevention.

epidemiology

Consequences Of Natural Perturbations In The Human Plasma Proteome

Proteins are the primary functional units of biology and the direct targets of most drugs, yet there is limited knowledge of the genetic factors determining inter-individual variation in protein levels. Here we reveal the genetic architecture of the human plasma proteome, testing 10.6 million DNA variants against levels of 2,994 proteins in 3,301 individuals. We identify 1,927 genetic associations with 1,478 proteins, a 4-fold increase on existing knowledge, including trans associations for 1,104 proteins. To understand consequences of perturbations in plasma protein levels, we introduce an approach that links naturally occurring genetic variation with biological, disease, and drug databases. We provide insights into pathogenesis by uncovering the molecular effects of disease-associated variants. We identify causal roles for protein biomarkers in disease through Mendelian randomization analysis. Our results reveal new drug targets, opportunities for matching existing drugs with new disease indications, and potential safety concerns for drugs under development.

genomics

Connecting genetic risk to disease endpoints through the human blood plasma proteome

Genome-wide association studies (GWAS) with intermediate phenotypes, like changes in metabolite and protein levels, provide functional evidence for mapping disease associations and translating them into clinical applications. However, although hundreds of genetic risk variants have been associated with complex disorders, the underlying molecular pathways often remain elusive. Associations with intermediate traits across multiple chromosome locations are key in establishing functional links between GWAS-identified risk-variants and disease endpoints. Here, we describe a GWAS performed with a highly multiplexed aptamer-based affinity proteomics platform. We quantified associations between protein level changes and gene variants in a German cohort and replicated this GWAS in an Arab/Asian cohort. We identified many independent, SNP-protein associations, which represent novel, inter-chromosomal links, related to autoimmune disorders, Alzheimer's disease, cardiovascular disease, cancer, and many other disease endpoints. We integrated this information into a genome-proteome network, and created an interactive web-tool for interrogations. Our results provide a basis for new approaches to pharmaceutical and diagnostic applications.

genetics

PopPAnTe: population and pedigree association testing for quantitative data

Family-based designs, from twin studies to isolated populations with their complex genealogical data, are a valuable resource for genetic studies of heritable molecular biomarkers. Existing software for family-based studies have mainly focused on facilitating association between response phenotypes and genetic markers, and no user-friendly tools are at present available to straightforwardly extend association studies in related samples to large datasets of generic quantitative data, as those generated by current -omics technologies.\n\nWe developed PopPAnTe, a user-friendly Java program, which evaluates the association of quantitative data in related samples. Additionally, Pop-PAnTe implements data pre and post processing, region based testing, and empirical assessment of associations.\n\nPopPAnTe is an integrated and flexible framework for pairwise association testing in related samples with a large number of predictors and response variables. It works either with family data of any size and complexity, or, when the genealogical information is unknown, it uses genetic similarity information between individuals as those inferred from genome-wide genetic data. It can therefore be particularly useful in facilitating usage of biobank data collections from population isolates when extensive genealogical information is missing.

bioinformatics