Search bioRxiv⌕ Search

Biology subjects

Wright, H. I. W.

Publications and source records attributed to Wright, H. I. W..

3 recordsLinked to original sources

Federated cross-biobank conditional analysis identifies LDL-C lowering effects of DNAJC13 haploinsufficiency and LDLR regulation

Whole genome sequencing in diverse population-scale biobanks offers new insights into the genetic architecture of complex traits from rare and non-coding variants. However, rare single variant and aggregate associations are often confounded by linkage disequilibrium and haplotype structure, resulting in large numbers of false-positive associations. Previous methods that rely on reference panels or linkage disequilibrium-matrices to determine conditional independence in meta-analyses do not scale to very rare variants, which may be observed in only one biobank and can exhibit long-range haplotypes. Here, we implement a federated approach to perform iterative conditional meta-analysis on individual-level genotype and phenotype data across biobanks while adhering to data sharing policies. We applied our methodology to a meta-analysis of LDL-C in 614,375 individuals from UK Biobank and All of Us, encompassing six genetic ancestry groups. After conditioning, only 4.3% of significantly associated rare single variants and 6.9% of aggregates remained statistically independent. The proportion of significant aggregates that remained independent after conditioning was higher for coding-based tests than non-coding. We further validate that our approach effectively suppresses false-positive associations using simulations centred on the LDLR locus. We identify allelic series of variants associated with reduced LDL-C, including loss-of-function variants in DNAJC13 and variants in the 3-prime untranslated region of LDLR. Our results highlight that federated conditioning can distinguish independent rare variant signals from linkage and haplotype structure artifacts in multi-ancestry meta-analyses across separate biobanks.

genetics↗

Genotype-level quality control substantially reduces error rates in population-scale whole-genome sequencing

Population-scale whole-genome sequencing data will contain many individual-level genotype errors, even after allele-level quality control (QC). We establish the need for genotype-level QC using UK Biobank (N=490,726) and All of Us v8 (N=414,830), where we remove up to 100 million ([~]9%) additional low-quality variants. We demonstrate reduced false positive rate in downstream genetic association studies, highlight the power of parent-offspring trios for QC, and illustrate the need for sex-specific X-chromosome filtering. We provide a QCed All of Us v8 dataset in plink-pgen format, and an efficient pipeline for QC and conversion from VCF to plink-pgen for UK Biobank.

genetics↗

Whole-genome sequencing analysis of anthropometric traits in 672,976 individuals reveals convergence between rare and common genetic associations

Genetic association studies have mostly focussed on common variants from genotyping arrays or rare protein-coding variants from exome sequencing. Here, we used whole-genome sequence (WGS) data in 672,976 individuals of diverse ancestry to evaluate the contribution and architecture of rare non-coding variants to three commonly studied anthropometric traits: height, body mass index (BMI) and waist-hip ratio adjusted for BMI (WHRadjBMI). Analysing 447,461 individuals in UK Biobank for discovery and 225,515 individuals in All of Us for replication, we identified 90 novel rare and low-frequency single variant associations. This includes two independent rare variants upstream of IGF2BP2 that both substantially reduce WHRadjBMI, but have distinct effects on other adiposity traits. We identified 135 coding variant aggregates, several of which were missed by exome sequencing studies. For example, UBR3 protein-truncating variants were associated with a 2.7kg/m2 increase in BMI. We additionally identified 51 non-coding variant aggregate associations, including in the 5UTR of FGF18 (a highly constrained gene with no previously reported coding associations) associated with up to 6cm effects on height. We show that 97% of rare variant associations occur near GWAS loci demonstrating convergence of rare and common variant associations. Finally, we show that ultra rare variants (MAF<0.01%) explain a small fraction of heritability (<10%) compared to common variants for these traits, that heritability is largely shared across ancestries, and that this heritability is concentrated at or near common variant loci. Our work demonstrates the importance of large-scale WGS for fully understanding the genetic architecture of complex traits.

genetics↗