Search bioRxivSearch

Biology subjects

Esko, T.

Publications and source records attributed to Esko, T..

14 recordsLinked to original sources

Polygenic prediction of breast cancer: comparison of genetic predictors and implications for screening

BackgroundPublished genetic risk scores for breast cancer (BC) so far have been based on a relatively small number of markers and are not necessarily using the full potential of large-scale Genome-Wide Association Studies. This study aims to identify an efficient polygenic predictor for BC based on best available evidence and to assess its potential for personalized risk prediction and screening strategies.\n\nMethodsFour different genetic risk scores (two already published and two newly developed) and their combinations (metaGRS) are compared in the subsets of two population-based biobank cohorts: the UK Biobank (UKBB, 3157 BC cases, 43,827 controls) and Estonian Biobank (EstBB, 317 prevalent and 308 incident BC cases in 32,557 women). In addition, correlations between different genetic risk scores and their associations with BC risk factors are studied in both cohorts.\n\nResultsThe metaGRS that combines two genetic risk scores (metaGRS2 - based on 75 and 898 Single Nucleotide Polymorphisms, respectively) has the strongest association with prevalent BC status in both cohorts. One standard deviation difference in the metaGRS2 corresponds to an Odds Ratio = 1.6 (95% CI 1.54 to 1.66, p = 9.7*10-135) in the UK Biobank and accounting for family history marginally attenuates the effect (Odds Ratio = 1.58, 95% CI 1.53 to 1.64, p = 9.1*10-129). In the EstBB cohort, the hazard ratio of incident BC for the women in the top 5% of the metaGRS2 compared to women in the lowest 50% is 4.2 (95% CI 2.8 to 6.2, p = 8.1*10-13). The different GRSs are only moderately correlated with each other and are associated with different known predictors of BC. The classification of genetic risk for the same individual may vary considerably depending on the chosen GRS.\n\nConclusionsWe have shown that metaGRS2 that combines on the effects of more than 900 SNPs provides best predictive ability for breast cancer in two different population-based cohorts. The strength of the effect of metaGRS2 indicates that the GRS could potentially be used to develop more efficient strategies for breast cancer screening for genotyped women.

genetics

The effect of X-linked dosage compensation on complex trait variation

Quantitative genetics theory predicts that X-chromosome dosage compensation between sexes will have a detectable effect on the amount of genetic and therefore phenotypic trait variances at associated loci in males and females. Here, we systematically examine the role of dosage compensation in complex trait variation in humans in 20 complex traits in a sample of more than 450,000 individuals from the UK Biobank and in 1,600 gene expression traits from a sample of 2,000 individuals as well as across-tissue gene expression from the GTEx resource. We find, on average, twice as much genetic variation for complex traits due to X-linked loci in males compared to females, consistent with a negligible effect of predicted escape from X-inactivation on complex trait variation across traits and also detect biologically relevant X-linked heterogeneity between the sexes for a number of complex traits.

genetics

Genomic analysis of diet composition finds novel loci and associations with health and lifestyle

We conducted genome-wide association study (GWAS) meta-analyses of relative caloric intake from fat, protein, carbohydrates and sugar in over 235,000 individuals. We identified 21 approximately independent lead SNPs. Relative protein intake exhibits the strongest relationships with poor health, including positive genetic associations with obesity, type 2 diabetes, and heart disease (rg {approx} 0.15 - 0.5). Relative carbohydrate and sugar intake have negative genetic correlations with waist circumference, waist-hip ratio, and neighborhood poverty (|rg| {approx} 0.1 - 0.3). Overall, our results show that the relative intake of each macronutrient has a distinct genetic architecture and pattern of genetic correlations suggestive of health implications beyond caloric content.

genetics

Genomic underpinnings of lifespan allow prediction and reveal basis in modern risks

We use a multi-stage genome-wide association of 1 million parental lifespans of genotyped subjects and data on mortality risk factors to validate previously unreplicated findings near CDKN2B-AS1, ATXN2/BRAP, FURIN/FES, ZW10, PSORS1C3, and 13q21.31, and identify and replicate novel findings near GADD45G, KCNK3, LDLR, POM121C, ZC3HC1, and ABO. We also validate previous findings near 5q33.3/EBF1 and FOXO3, whilst finding contradictory evidence at other loci. Gene set and tissue-specific analyses show that expression in foetal brain cells and adult dorsolateral prefrontal cortex is enriched for lifespan variation, as are gene pathways involving lipid proteins and homeostasis, vesicle-mediated transport, and synaptic function. Individual genetic variants that increase dementia, cardiovascular disease, and lung cancer -but not other cancers-explain the most variance, possibly reflecting modern susceptibilities, whilst cancer may act through many rare variants, or the environment. Resultant polygenic scores predict a mean lifespan difference of around five years of life across the deciles.

genomics

PROTEIN-CODING VARIANTS IMPLICATE NOVEL GENES RELATED TO LIPID HOMEOSTASIS CONTRIBUTING TO BODY FAT DISTRIBUTION

Body fat distribution is a heritable risk factor for a range of adverse health consequences, including hyperlipidemia and type 2 diabetes. To identify protein-coding variants associated with body fat distribution, assessed by waist-to-hip ratio adjusted for body mass index, we analyzed 228,985 predicted coding and splice site variants available on exome arrays in up to 344,369 individuals from five major ancestries for discovery and 132,177 independent European-ancestry individuals for validation. We identified 15 common (minor allele frequency, MAF[&ge;]5%) and 9 low frequency or rare (MAF<5%) coding variants that have not been reported previously. Pathway/gene set enrichment analyses of all associated variants highlight lipid particle, adiponectin level, abnormal white adipose tissue physiology, and bone development and morphology as processes affecting fat distribution and body shape. Furthermore, the cross-trait associations and the analyses of variant and gene function highlight a strong connection to lipids, cardiovascular traits, and type 2 diabetes. In functional follow-up analyses, specifically in Drosophila RNAi-knockdown crosses, we observed a significant increase in the total body triglyceride levels for two genes (DNAH10 and PLXND1). By examining variants often poorly tagged or entirely missed by genome-wide association studies, we implicate novel genes in fat distribution, stressing the importance of interrogating low-frequency and protein-coding variants.

genetics

A study of analytical strategies to include X-chromosome in variance heterogeneity analysis: evidence for trait-specific polygenic variance structure

Genotype-stratified variance of a quantitative trait could differ in the presence of gene-gene or gene-environment interactions. Genetic markers associated with phenotypic variance are thus considered promising candidates for follow-up interaction or joint location-scale analyses. However, as in studies of main effects, the X-chromosome is routinely excluded from whole-genome scans due to analytical challenges. Specifically, as males carry only one copy of the X-chromosome, the inherent sex-genotype dependency could bias the trait-genotype association, through sexual dimorphism in quantitative traits with sex-specific means or variances. Here we investigate phenotypic variance heterogeneity associated with X-chromosome SNPs and propose valid and powerful strategies. Among those, a generalized Levenes test has adequate power and remains robust to sexual dimorphism. An alternative approach is sex-stratified analysis but at the cost of slightly reduced power and modeling flexibility. We applied both methods to an Estonian study of gene expression quantitative trait loci (eQTL; n=841), and two complex trait studies of height, hip and waist circumferences, and body mass index from multi-ethnic study of atherosclerosis (MESA; n=2,073) and UK Biobank (UKB; n=327,393). Consistent with previous eQTL findings on mean, we found some but no conclusive evidence for cis regulators being enriched for variance association. SNP rs2681646 is associated with variance of waist circumference (p=9.5E-07) at X-chromosome-wide significance in UKB, with a suggestive female-specific effect in MESA (p=0.048). Collectively, an enrichment analysis using permutated UKB (p<1/10) and MESA (p<1/100) datasets, suggests a possible polygenic structure for the variance of human height.

genetics

Deep coverage whole genome sequences and plasma lipoprotein(a) in individuals of European and African ancestries

Lipoprotein(a), Lp(a), is a modified low-density lipoprotein particle where apolipoprotein(a) (protein product of the LPA gene) is covalently attached to apolipoprotein B. Lp(a) is a highly heritable, causal risk factor for cardiovascular diseases and varies in concentrations across ancestries. To comprehensively delineate the inherited basis for plasma Lp(a), we performed deep-coverage whole genome sequencing in 8,392 individuals of European and African American ancestries. Through whole genome variant discovery and direct genotyping of all structural variants overlapping LPA, we quantified the 5.5kb kringle IV-2 copy number (KIV2-CN), a known LPA structural polymorphism, and developed a model for its imputation. Through common variant analysis, we discovered a novel locus (SORT1) associated with Lp(a)-cholesterol, and also genetic modifiers of KIV2-CN. Furthermore, in contrast to previous GWAS studies, we explain most of the heritability of Lp(a), observing Lp(a) to be 85% heritable among African Americans and 75% among Europeans, yet with notable inter-ethnic heterogeneity. Through analyses of aggregates of rare coding and non-coding variants with Lp(a)-cholesterol, we found the only genome-wide significant signal to be at a non-coding SLC22A3 intronic window also previously described to be associated with Lp(a); however, this association was mitigated by adjustment with KIV2-CN. Finally, using an additional imputation dataset (N=27,344), we performed Mendelian randomization of LPA variant classes, finding that genetically regulated Lp(a) is more strongly associated with incident cardiovascular diseases than directly measured Lp(a), and is significantly associated with measures of subclinical atherosclerosis in African Americans.

genomics

Deep-coverage whole genome sequences and blood lipids among 16,324 individuals

Deep-coverage whole genome sequencing at the population level is now feasible and offers potential advantages for locus discovery, particularly in the analysis rare mutations in non-coding regions. Here, we performed whole genome sequencing in 16,324 participants from four ancestries at mean depth >29X and analyzed correlations of genotypes with four quantitative traits - plasma levels of total cholesterol, low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol, and triglycerides. We conducted a discovery analysis including common or rare variants in coding as well as non-coding regions and developed a framework to interpret genome sequence for dyslipidemia risk. Common variant association yielded loci previously described with the exception of a few variants not captured earlier by arrays or imputation. In coding sequence, rare variant association yielded known Mendelian dyslipidemia genes and, in non-coding sequence, we detected no rare variant association signals after application of four approaches to aggregate variants in non-coding regions. We developed a new, genome-wide polygenic score for LDL-C and observed that a high polygenic score conferred similar effect size to a monogenic mutation (~30 mg/dl higher LDL-C for each); however, among those with extremely high LDL-C, a high polygenic score was considerably more prevalent than a monogenic mutation (23% versus 2% of participants, respectively).

genomics

PAIRUP-MS: Pathway Analysis and Imputation to Relate Unknowns in Profiles from Mass Spectrometry-based metabolite data

Metabolomics is a powerful approach for discovering biomarkers and metabolic quantitative trait loci. While untargeted profiling methods can measure up to thousands of metabolite signals in a single experiment, many signals cannot be readily identified as known metabolites or compared across datasets, making it difficult to infer biology and to conduct well-powered meta-analyses across studies. To deal with these challenges, we developed a suite of computational methods, PAIRUP-MS, to match metabolite signals across mass spectrometry-based profiling datasets using an imputation-based approach and to generate pathway annotations for these signals. We performed meta and pathway analyses for both known and unknown signals in multiple datasets and then validated the results using genetic associations. Finally, we applied the methods to detect metabolite signals and pathways associated with body mass index, demonstrating that our framework is useful for analyzing unknown signals in a robust and biologically meaningful manner and for improving the power of untargeted metabolomics studies.

bioinformatics

Haplotype sharing provides insights into fine-scale population history and disease in Finland

Finland provides unique opportunities to investigate population and medical genomics because of its adoption of unified national electronic health records, detailed historical and birth records, and serial population bottlenecks. We assemble a comprehensive view of recent population history ([&le;]100 generations), the timespan during which most rare disease-causing alleles arose, by comparing pairwise haplotype sharing from 43,254 Finns to geographically and linguistically adjacent countries with different population histories, including 16,060 Swedes, Estonians, Russians, and Hungarians. We find much more extensive sharing in Finns, with at least one [&ge;] 5 cM tract on average between pairs of unrelated individuals. By coupling haplotype sharing with fine-scale birth records from over 25,000 individuals, we find that while haplotype sharing broadly decays with geographical distance, there are pockets of excess haplotype sharing; individuals from northeast Finland share several-fold more of their genome in identity-by-descent (IBD) segments than individuals from southwest regions containing the major cities of Helsinki and Turku. We estimate recent effective population size changes over time across regions of Finland and find significant differences between the Early and Late Settlement Regions as expected; however, our results indicate more continuous gene flow than previously indicated as Finns migrated towards the northernmost Lapland region. Lastly, we show that haplotype sharing is locally enriched among pairs of individuals sharing rare alleles by an order of magnitude, especially among pairs sharing rare disease causing variants. Our work provides a general framework for using haplotype sharing to reconstruct an integrative view of recent population history and gain insight into the evolutionary origins of rare variants contributing to disease.

genetics

Genetic analysis of over one million people identifies 535 novel loci for blood pressure.

High blood pressure is the foremost heritable global risk factor for cardiovascular disease. We report the largest genetic association study of blood pressure traits to date (systolic, diastolic, pulse pressure) in over one million people of European ancestry. We identify 535 novel blood pressure loci that not only offer new biological insights into blood pressure regulation but also reveal shared loci influencing lifestyle exposures. Our findings offer the potential for a precision medicine strategy for future cardiovascular disease prevention.

genetics

Genome-wide Association Study of Clinical Features in the Schizophrenia Psychiatric Genomics Consortium: Confirmation of Polygenic Effect on Negative Symptoms

Schizophrenia is a clinically heterogeneous disorder. Proposed revisions in DSM - 5 included dimensional measurement of different symptom domains. We sought to identify common genetic variants influencing these dimensions, and confirm a previous association between polygenic risk of schizophrenia and the severity of negative symptoms. The Psychiatric Genomics Consortium study of schizophrenia comprised 8,432 cases of European ancestry with available clinical phenotype data. Symptoms averaged over the course of illness were assessed using the OPCRIT, PANSS, LDPS, SCAN, SCID, and CASH. Factor analyses of each constituent PGC study identified positive, negative, manic, and depressive symptom dimensions. We examined the relationship between the resultant symptom dimensions and aggregate polygenic risk scores indexing risk of schizophrenia. We performed genome - wide association study (GWAS) of each quantitative traits using linear regression and adjusting for significant effects of sex and ancestry. The negative symptom factor was significantly associated with polygene risk scores for schizophrenia, confirming a previous, suggestive finding by our group in a smaller sample, though explaining only a small fraction of the variance. In subsequent GWAS, we observed the strongest evidence of association for the positive and negative symptom factors, with SNPs in RFX8 on 2q11.2 (P = 6.27x10-8) and upstream of WDR72 / UNC13C on 15q21.3 (P = 7.59x10-8), respectively. We report evidence of association of novel modifier loci for schizophrenia, though no single locus attained established genome - wide significance criteria. As this may have been due to insufficient statistical power, follow - up in additional samples is warranted. Importantly, we replicated our previous finding that polygenic risk explains at least some of the variance in negative symptoms, a core illness dimension.

genetics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics

Constraints on eQTL fine mapping in the presence of multi-site local regulation of gene expression

Expression QTL (eQTL) detection has emerged as an important tool for unravelling of the relationship between genetic risk factors and disease or clinical phenotypes. Most studies use single marker linear regression to discover primary signals, followed by sequential conditional modeling to detect secondary genetic variants affecting gene expression. However, this approach assumes that functional variants are sparsely distributed and that close linkage between them has little impact on estimation of their precise location and magnitude of effects. In this study, we address the prevalence of secondary signals and bias in estimation of their effects by performing multi-site linear regression on two large human cohort peripheral blood gene expression datasets (each greater than 2,500 samples) with accompanying whole genome genotypes, namely the CAGE compendium of Illumina microarray studies, and the Framingham Heart Study Affymetrix data. Stepwise conditional modeling demonstrates that multiple eQTL signals are present for ~40% of over 3500 eGenes in both datasets, and the number of loci with additional signals reduces by approximately two-thirds with each conditioning step. However, the concordance of specific signals between the two studies is only ~30%, indicating that expression profiling platform is a large source of variance in effect estimation. Furthermore, a series of simulation studies imply that in the presence of multi-site regulation, up to 10% of the secondary signals could be artefacts of incomplete tagging, and at least 5% but up to one quarter of credible intervals may not even include the causal site, which is thus mis-localized. Joint multi-site effect estimation recalibrates effect size estimates by just a small amount on average. Presumably similar conclusions apply to most types of quantitative trait. Given the strong empirical evidence that gene expression is commonly regulated by more than one variant, we conclude that the fine-mapping of causal variants needs to be adjusted for multi-site influences, as conditional estimates can be highly biased by interference among linked sites.

genetics