Search bioRxivSearch

Biology subjects

Hilary Kiyo Finucane

Publications and source records attributed to Hilary Kiyo Finucane.

6 recordsLinked to original sources

Correcting subtle stratification in summary association statistics

Population stratification is a well-documented confounder in GWASes, and is often addressed by including principal component (PC) covariates computed from common SNPs (SNP-PCs). In our analyses of summary statistics from 36 GWASes (mean n=88k), including 20 GWASes using 23andMe data that included SNP-PC covariates, we observed a significantly inflated LD score regression (LDSC) intercept for several traits--suggesting that residual stratification remains a concern, even when SNPPC covariates are included.\n\nHere we propose a new method, PC loading regression, to correct for stratification in summary statistics by leveraging SNP loadings for PCs computed in a large reference panel. In addition to SNP-PCs, the method can be applied to haploSNP-PCs, i.e. PCs computed from a larger number of rare haplotype variants that better capture subtle structure. Using simulations based on real genotypes from 54,000 individuals of diverse European ancestry from the Genetic Epidemiology Research on Adult Health and Aging (GERA) cohort, we show that PC loading regression effectively corrects for stratification along top PCs.\n\nWe applied PC loading regression to several traits with inflated LDSC intercepts. Correcting for the top four SNP-PCs in GERA data, we observe a significant reduction in LDSC intercept height summary statistics from the Genetic Investigation of ANthropometric Traits (GIANT) consortium, but not for 23andMe summary statistics, which already included SNP-PC covariates. However, when correcting for additional haploSNP-PCs in 23andMe GWASes, inflation in the LDSC intercept was eliminated for eye color, hair color, and skin color and substantially reduced for height (1.41 to 1.16; n=430k). Correcting for haploSNP-PCs in GIANT height summary statistics eliminated inflation in the LDSC intercept (from 1.35 to 1.00; n=250k), eliminating 27 significant association signals including one at the LCT locus, which is highly differentiated among European populations and widely known to produce spurious signals. Overall, our results suggest that uncorrected population stratification is a concern in GWASes of large sample size and that PC loading regression can correct for this stratification.

Genetics

Analysis of shared heritability in common disorders of the brain

Disorders of the brain exhibit considerable epidemiological comorbidity and frequently share symptoms, provoking debate about the extent of their etiologic overlap. We quantified the genetic sharing of 25 brain disorders based on summary statistics from genome-wide association studies of 215,683 patients and 657,164 controls, and their relationship to 17 phenotypes from 1,191,588 individuals. Psychiatric disorders show substantial sharing of common variant risk, while neurological disorders appear more distinct from one another. We observe limited evidence of sharing between neurological and psychiatric disorders, but do identify robust sharing between disorders and several cognitive measures, as well as disorders and personality types. We also performed extensive simulations to explore how power, diagnostic misclassification and phenotypic heterogeneity affect genetic correlations. These results highlight the importance of common genetic variation as a source of risk for brain disorders and the value of heritability-based methods in understanding their etiology.

Genetics

Subtle stratification confounds estimates of heritability from rare variants

Genome-wide significant associations generally explain only a small proportion of the narrow-sense heritability of complex disease (h2). While considerably more heritability is explained by all genotyped SNPs (hg2), for most traits, much heritability remains missing (hg2 < h2). Rare variants, poorly tagged by genotyped SNPs, are a major potential source of the gap between hg2 and h2. Recent efforts to assess the contribution of both sequenced and imputed rare variants to phenotypes suggest that substantial heritability may lie in these variants. Here we analyze sequenced SNPs, imputed SNPs and haploSNPs-- haplotype variants constructed from within a sample, without using a reference panel-- and show that studies of heritability from these variants may be strongly confounded by subtle population stratification. For example, when meta-analyzing heritability estimates from 22 randomly ascertained case-control traits from the GERA cohort, we observe a statistically significant increase in heritability explained by imputed SNPs even after correcting for principal components (PCs) from genotyped (or imputed) SNPs. However, this increase is eliminated when correcting for stratification using PCs from a larger number of haploSNPs. We note that subtle stratification may also impact estimates of heritability from array SNPs, although we find that this is generally a less severe problem. Overall, our results suggest that estimating the heritability explained by rare variants for case-control traits requires exquisite control for population stratification, but current methods may not provide this level of control.

Genetics

Functional partitioning of local and distal gene expression regulation in multiple human tissues

Studies of the genetics of gene expression have served as a key tool for linking genetic variants to phenotypes. Large-scale eQTL mapping studies have identified a large number of local eQTLs, but the molecular mechanism of how genetic variants regulate expression is still unclear, particularly for distal eQTLs, which these studies are not well-powered to detect. In this study, we use a heritability partitioning approach to dissect the functional components of gene regulation. We make use of an existing method, stratified LD score regression, that leverages all variants (not just those that pass stringent significance thresholds) to partition heritability across functional categories, and we extend this method to partition local and distal gene expression heritability in 15 human tissues. The top enriched functional categories in local regulation of peripheral blood gene expression included super enhancers (5.18x), coding regions (3.73x), conserved regions (2.33x) and four histone marks (p<3x10-7 for all enrichments); local enrichments were similar across the 15 tissues. We also observed substantial enrichments for distal regulation of peripheral blood gene expression: super enhancers (1.91x), coding regions (4.47x), conserved regions (4.51x) and two histone marks (p<3x10-7 for all enrichments). Analyses of the genetic correlation of gene expression across tissues showed that local gene expression regulation is largely shared across tissues, but distal gene expression regulation is highly tissue-specific. Our results elucidate the functional components of the genetic architecture of local and distal gene expression regulation.

Genomics

Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores

Polygenic risk scores have shown great promise in predicting complex disease risk, and will become more accurate as training sample sizes increase. The standard approach for calculating risk scores involves LD-pruning markers and applying a P-value threshold to association statistics, but this discards information and may reduce predictive accuracy. We introduce a new method, LDpred, which infers the posterior mean causal effect size of each marker using a prior on effect sizes and LD information from an external reference panel. Theory and simulations show that LDpred outperforms the pruning/thresholding approach, particularly at large sample sizes. Accordingly, prediction R2 increased from 20.1% to 25.3% in a large schizophrenia data set and from 9.8% to 12.0% in a large multiple sclerosis data set. A similar relative improvement in accuracy was observed for three additional large disease data sets and when predicting in non-European schizophrenia samples. The advantage of LDpred over existing methods will grow as sample sizes increase.

Bioinformatics

Partitioning heritability by functional category using GWAS summary statistics

Recent work has demonstrated that some functional categories of the genome contribute disproportionately to the heritability of complex diseases. Here, we analyze a broad set of functional elements, including cell-type-specific elements, to estimate their polygenic contributions to heritability in genome-wide association studies (GWAS) of 17 complex diseases and traits spanning a total of 1.3 million phenotype measurements. To enable this analysis, we introduce a new method for partitioning heritability from GWAS summary statistics while controlling for linked markers. This new method is computationally tractable at very large sample sizes, and leverages genome-wide information. Our results include a large enrichment of heritability in conserved regions across many traits; a very large immunological disease-specific enrichment of heritability in FANTOM5 enhancers; and many cell-type-specific enrichments including significant enrichment of central nervous system cell types in body mass index, age at menarche, educational attainment, and smoking behavior. These results demonstrate that GWAS can aid in understanding the biological basis of disease and provide direction for functional follow-up.

Genetics