Search bioRxivSearch

Biology subjects

Yengo, L.

Publications and source records attributed to Yengo, L..

9 recordsLinked to original sources

Machine Learning in Multi-Omics Data to Assess Longitudinal Predictors of Glycaemic Trait Levels

Type 2 diabetes (T2D) is a global health burden that will benefit from personalised risk prediction and targeted prevention programmes. Omics data have enabled more detailed risk prediction; however, most studies have focussed on directly on the ability of DNA variants predicting T2D onset with less attention given to epigenetic regulation and glycaemic trait variability. By applying machine learning to the longitudinal Northern Finland Birth Cohort 1966 (NFBC 1966) at 31 (T1) and 46 (T2) years old, we predicted fasting glucose (FG) and insulin (FI), glycated haemoglobin (HbA1c) and 2-hour glucose and insulin from oral glucose tolerance test (2hGlu, 2hIns) at T2 in 513 individuals from 1,001 variables at T1 and T2, including anthropometric, metabolic, metabolomic and epigenetic variables. We further tested whether the information obtained by the machine learning models in NFBC could be used to predict glycaemic traits in the independent French study with 48 matching predictors (DESIR, N=769, age range 30-65 years at recruitment, interval between data collections: 9 years). In this study, FG and FI were best predicted, with average R2 values of 0.38 and 0.53. Sex, branched-chain and aromatic amino acids, HDL-cholesterol, glycerol, ketone bodies, blood pressure at T2 and measurements of adiposity at T1, as well as multiple methylation marks at both time points were amongst the top predictors. In the validation analysis, we reached R2 values of 0.41/0.55 for FG/FI when trained and tested in NFBC1966 and 0.17/0.30 when trained in NFBC1966 and tested in DESIR. We identified clinically relevant sets of predictors from a large multi-omics dataset and highlighted the potential of methylation markers and longitudinal changes in prediction.

genomics

Expectation of the intercept from bivariate LD score regression in the presence of population stratification

Linkage disequilibrium (LD) score regression is an increasingly popular method used to quantify the level of confounding in genome-wide association studies (GWAS) or to estimate heritability and genetic correlation between traits. When applied to a pair of GWAS, the LD score regression (LDSC) methodology produces a statistic, referred to as the bivariate LDSC intercept, which deviation from 0 is classically interpreted as an indication of sample overlap between the two GWAS. Here we propose an extension of the theory underlying the bivariate LDSC methodology, which accounts for population stratification within and between GWAS. Our extended theory predicts an inflation of the bivariate LDSC intercept when sample sizes and heritability are large, even in the absence of sample overlap. We illustrate our theoretical results with simulations based on actual SNP genotypes and we propose a re-interpretation of previously published results in the light of our extended theory.

genetics

Meta-analysis of genome-wide association studies for body fat distribution in 694,649 individuals of European ancestry

One in four adults worldwide are either overweight or obese. Epidemiological studies indicate that the location and distribution of excess fat, rather than general adiposity, is most informative for predicting risk of obesity sequellae, including cardiometabolic disease and cancer. We performed a genome-wide association study meta-analysis of body fat distribution, measured by waist-to-hip ratio adjusted for BMI (WHRadjBMI), and identified 463 signals in 346 loci. Heritability and variant effects were generally stronger in women than men, and we found approximately one-third of all signals to be sexually dimorphic. The 5% of individuals carrying the most WHRadjBMI-increasing alleles were 1.62 times more likely than the bottom 5% to have a WHR above the thresholds used for metabolic syndrome. These data, made publicly available, will inform the biology of body fat distribution and its relationship with disease.

genetics

Imprint of Assortative Mating on the Human Genome

Non-random mate-choice with respect to complex traits is widely observed in humans, but whether this reflects true phenotypic assortment, environment (social homogamy) or convergence after choosing a partner is not known. Understanding the causes of mate choice is important, because assortative mating (AM) if based upon heritable traits, has genetic and evolutionary consequences. AM is predicted under Fishers classical theory1 to induce a signature in the genome at trait-associated loci that can be detected and quantified. Here, we develop and apply a method to quantify AM on a specific trait by estimating the correlation ({theta}) between genetic predictors of the trait from SNPs on odd versus even chromosomes. We show by theory and simulation that the effect of AM can be distinguished from population stratification. We applied this approach to 32 complex traits and diseases using SNP data from [~]400,000 unrelated individuals of European ancestry. We found significant evidence of AM for height ({theta}=3.2%) and educational attainment ({theta}=2.7%), both consistent with theoretical predictions. Overall, our results imply that AM involves multiple traits, affects the genomic architecture of loci that are associated with these traits and that the consequence of mate choice can be detected from a random sample of genomes.

genetics

Novel susceptibility loci and genetic regulation mechanisms for type 2 diabetes

We conducted a meta-analysis of genome-wide association studies (GWAS) with [~]16 million genotyped/imputed genetic variants in 62,892 type 2 diabetes (T2D) cases and 596,424 controls of European ancestry. We identified 139 common and 4 rare (minor allele frequency < 0.01) variants associated with T2D, 42 of which (39 common and 3 rare variants) were independent of the known variants. Integration of the gene expression data from blood (n = 14,115 and 2,765) and other T2D-relevant tissues (n = up to 385) with the GWAS results identified 33 putative functional genes for T2D, three of which were targeted by approved drugs. A further integration of DNA methylation (n = 1,980) and epigenomic annotations data highlighted three putative T2D genes (CAMK1D, TP53INP1 and ATP5G1) with plausible regulatory mechanisms whereby a genetic variant exerts an effect on T2D through epigenetic regulation of gene expression. We further found evidence that the T2D-associated loci have been under purifying selection.

genetics

Identifying gene targets for brain-related traits using transcriptomic and methylomic data from blood

Understanding the difference in genetic regulation of gene expression between brain and blood is important for discovering genes associated with brain-related traits and disorders. Here, we estimate the correlation of genetic effects at the top associated cis-expression (cis-eQTLs or cis-mQTLs) between brain and blood for genes expressed (or CpG sites methylated) in both tissues, while accounting for errors in their estimated effects (rb). Using publicly available data (n = 72 to l,366), we find that the genetic effects of cis-eQTLs (PeQTL < 5x10-8) or mQTLs (PmQTL < 1x10-10) are highly correlated between independent brain and blood samples ([Formula] with SE = 0.015 for cis-eQTL and [Formula] with SE = 0.006 for cis-mQTLs). Using meta-analyzed brain eQTL/mQTL data (n = 526 to 1,194), we identify 61 genes and 167 DNA methylation (DNAm) sites associated with 4 brain-related traits and disorders. Most of these associations are a subset of the discoveries (97 genes and 295 DNAm sites) using data from blood with larger sample sizes (n = l,980 to 14,115). We further find that cis-eQTLs with tissue-specific effects are approximately uniformly distributed across all the functional annotation categories, and that mean difference in gene expression level between brain and blood is almost independent of the difference in the corresponding cis-eQTL effect. Our results demonstrate the gain of power in gene discovery for brain-related phenotypes using blood cis-eQTL or cis-mQTL data with large sample sizes.

genetics

Meta-analysis of genome-wide association studies for height and body mass index in ~700,000 individuals of European ancestry

Genome-wide association studies (GWAS) stand as powerful experimental designs for identifying DNA variants associated with complex traits and diseases. In the past decade, both the number of such studies and their sample sizes have increased dramatically. Recent GWAS of height and body mass index (BMI) in [~]250,000 European participants have led to the discovery of [~]700 and [~]100 nearly independent SNPs associated with these traits, respectively. Here we combine summary statistics from those two studies with GWAS of height and BMI performed in [~]450,000 UK Biobank participants of European ancestry. Overall, our combined GWAS meta-analysis reaches N[~]700,000 individuals and substantially increases the number of GWAS signals associated with these traits. We identified 3,290 and 716 near-independent SNPs associated with height and BMI, respectively (at a revised genome-wide significance threshold of p<1 x 10-8), including 1,185 height-associated SNPs and 554 BMI-associated SNPs located within loci not previously identified by these two GWAS. The genome-wide significant SNPs explain [~]24.6% of the variance of height and [~]5% of the variance of BMI in an independent sample from the Health and Retirement Study (HRS). Correlations between polygenic scores based upon these SNPs with actual height and BMI in HRS participants were 0.44 and 0.20, respectively. From analyses of integrating GWAS and eQTL data by Summary-data based Mendelian Randomization (SMR), we identified an enrichment of eQTLs amongst lead height and BMI signals, prioritisting 684 and 134 genes, respectively. Our study demonstrates that, as previously predicted, increasing GWAS sample sizes continues to deliver, by discovery of new loci, increasing prediction accuracy and providing additional data to achieve deeper insight into complex trait biology. All summary statistics are made available for follow up studies.

genetics

Fine-mapping of an expanded set of type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps

We aggregated genome-wide genotyping data from 32 European-descent GWAS (74,124 T2D cases, 824,006 controls) imputed to high-density reference panels of >30,000 sequenced haplotypes. Analysis of {small tilde}27M variants ({small tilde}21M with minor allele frequency [MAF]<5%), identified 243 genome-wide significant loci (p<5x10-8; MAF 0.02%-50%; odds ratio [OR] 1.04-8.05), 135 not previously-implicated in T2D-predisposition. Conditional analyses revealed 160 additional distinct association signals (p<10-5) within the identified loci. The combined set of 403 T2D-risk signals includes 56 low-frequency (0.5%[&le;]MAF<5%) and 24 rare (MAF<0.5%) index SNPs at 60 loci, including 14 with estimated allelic OR>2. Forty-one of the signals displayed effect-size heterogeneity between BMI-unadjusted and adjusted analyses. Increased sample size and improved imputation led to substantially more precise localisation of causal variants than previously attained: at 51 signals, the lead variant after fine-mapping accounted for >80% posterior probability of association (PPA) and at 18 of these, PPA exceeded 99%. Integration with islet regulatory annotations enriched for T2D association further reduced median credible set size (from 42 variants to 32) and extended the number of index variants with PPA>80% to 73. Although most signals mapped to regulatory sequence, we identified 18 genes as human validated therapeutic targets through coding variants that are causal for disease. Genome wide chip heritability accounted for 18% of T2D-risk, and individuals in the 2.5% extremes of a polygenic risk score generated from the GWAS data differed >9-fold in risk. Our observations highlight how increases in sample size and variant diversity deliver enhanced discovery and single-variant resolution of causal T2D-risk alleles, and the consequent impact on mechanistic insights and clinical translation.

genomics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics