Search bioRxivSearch

Biology subjects

Finucane, H. K.

Publications and source records attributed to Finucane, H. K..

4 recordsLinked to original sources

Reconciling S-LDSC and LDAK functional enrichment estimates

Recent work has highlighted the importance of accounting for linkage disequilibrium (LD)-dependent genetic architectures in analyses of heritability, motivating the development of the baseline-LD model used by stratified LD score regression (S-LDSC) and the LDAK model. Although both models include LD-dependent effects, they produce very different estimates of functional enrichment (with larger estimates using the baseline-LD model), leading to different interpretations of the functional architecture of complex traits. Here, we perform formal model comparisons and empirical analyses to reconcile these findings. First, by performing model comparisons using a likelihood approach, we determined that the baseline-LD model attains likelihoods across 16 UK Biobank traits that are substantially higher than the LDAK model. Second, we determined that S-LDSC using a combined model (unlike methods that use the LDAK or baseline-LD models) produces robust enrichment estimates in simulations under both the LDAK and baseline-LD models, validating the combined model as a gold standard. Third, in analyses of 16 UK Biobank traits, we determined that enrichment estimates obtained by S-LDSC using the combined model were nearly identical to those obtained by S-LDSC using the baseline-LD model (concordance correlation coefficient{rho} c = 0.99), but were larger than those obtained using LDAK ({rho}c = 0.54). Notably, LDAK enrichment estimates were much higher for a non-default version of LDAK that models SNPs in perfect LD differently by assigning non-zero weights to all SNPs. Our results support the use of the baseline-LD model and confirm the existence of functional annotations that are highly enriched for complex trait heritability.

genetics

Interrogation of human hematopoiesis at single-cell and single-variant resolution

Incomplete annotation of cell-to-cell state variance and widespread linkage disequilibrium in the human genome represent significant challenges to elucidating mechanisms of trait-associated genetic variation. Here, using data from the UK Biobank, we perform genetic fine-mapping for 16 blood cell traits to quantify posterior probabilities of association while allowing for multiple independent signals per region. We observe an enrichment of fine-mapped variants in accessible chromatin of lineage-committed hematopoietic progenitor cells. Further, we develop a novel analytic framework that identifies \"core gene\" cell type enrichments and show that this approach uniquely resolves relevant cell types within closely related populations. Applying our approach to single cell chromatin accessibility data, we discover significant heterogeneity within classically defined multipotential progenitor populations. Finally, using several lines of empirical evidence, we identify relevant cell types, predict target genes, and propose putative causal mechanisms for fine-mapped variants. In total, our study provides an analytic framework for single-variant and single-cell analyses to elucidate putative causal variants and cell types from GWAS and high-resolution epigenomic assays.

genetics

Detecting genome-wide directional effects of transcription factor binding on polygenic disease risk

Biological interpretation of GWAS data frequently involves analyzing unsigned genomic annotations comprising SNPs involved in a biological process and assessing enrichment for disease signal. However, it is often possible to generate signed annotations quantifying whether each SNP allele promotes or hinders a biological process, e.g., binding of a transcription factor (TF). Directional effects of such annotations on disease risk enable stronger statements about causal mechanisms of disease than enrichments of corresponding unsigned annotations. Here we introduce a new method, signed LD profile regression, for detecting such directional effects using GWAS summary statistics, and we apply the method using 382 signed annotations reflecting predicted TF binding. We show via theory and simulations that our method is well-powered and is well-calibrated even when TF binding sites co-localize with other enriched regulatory elements, which can confound unsigned enrichment methods. We further validate our method by showing that it recovers known transcriptional regulators when applied to molecular QTL in blood. We then apply our method to eQTL in 48 GTEx tissues, identifying 651 distinct TF-tissue expression associations at per-tissue FDR < 5%, including 30 associations with robust evidence of tissue specificity. Finally, we apply our method to 46 diseases and complex traits (average N = 289,617) and identify 77 annotation-trait associations at per-trait FDR < 5% representing 12 independent TF-trait associations, and we conduct gene-set enrichment analyses to characterize the underlying transcriptional programs. Our results implicate new causal disease genes (including causal genes at known GWAS loci), and in some cases suggest a detailed mechanism for a causal genes effect on disease. Our method provides a new way to leverage functional data to draw inferences about disease etiology.

genetics

Estimating the proportion of disease heritability mediated by gene expression levels

Disease risk variants identified by GWAS are predominantly noncoding, suggesting that gene regulation plays an important role. eQTL studies in unaffected individuals are often used to link disease-associated variants with the genes they regulate, relying on the hypothesis that noncoding regulatory effects are mediated by steady-state expression levels. To test this hypothesis, we developed a method to estimate the proportion of disease heritability mediated by the cis-genetic component of assayed gene expression levels. The method, gene expression co-score regression (GECS regression), relies on the idea that, for a gene whose expression level affects a phenotype, SNPs with similar effects on the expression of that gene will have similar phenotypic effects. In order to distinguish directional effects mediated by gene expression from non-directional pleiotropic or tagging effects, GECS regression operates on pairs of cis SNPs in linkage equilibrium, regressing pairwise products of disease effect sizes on products of cis-eQTL effect sizes. We verified that GECS regression produces robust estimates of mediated effects in simulations. We applied the method to eQTL data in 44 tissues from the GTEx consortium (average NeQTL = 158 samples) in conjunction with GWAS summary statistics for 30 diseases and complex traits (average NGWAS = 88K) with low pairwise genetic correlation, estimating the proportion of SNP-heritability mediated by the cis-genetic component of assayed gene expression in the union of the 44 tissues. The mean estimate was 0.21 (s.e. = 0.01) across 30 traits, with a significantly positive estimate (p < 0.001) for every trait. Thus, assayed gene expression in bulk tissues mediates a statistically significant but modest proportion of disease heritability, motivating the development of additional assays to capture regulatory effects and the use of our method to estimate how much disease heritability they mediate.

genetics