Search bioRxiv⌕ Search

Biology subjects

Hawkes, G.

Publications and source records attributed to Hawkes, G..

3 recordsLinked to original sources

Whole genome association testing in 333,100 individuals across three biobanks identifies rare non-coding single variant and genomic aggregate associations with height

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N=200,003), TOPMed (N=87,652) and All of Us (N=45,445). We performed rare (<0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal, intergenic and deep-intronic annotation. We observed 29 independent variants associated with height at P < 6 x 10-10 after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed a novel approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

genetics↗

Whole genome sequencing analysis identifies rare, large-effect non-coding variants and regions associated with circulating protein levels

The role of non-coding rare variation in common phenotypes is largely unknown, due to a lack of whole-genome sequence data, and the difficulty of categorising non-coding variants into biologically meaningful regulatory units. To begin addressing these challenges, we performed a cis association analysis using whole-genome sequence data, consisting of 391 million variants and 1,450 circulating protein levels in [~]20,000 UK Biobank participants. We identified 777 independent rare non-coding single variants associated with circulating protein levels (P<1x10-9), after conditioning on protein-coding and common associated variants. Rare non-coding aggregate testing identified 108 conditionally independent regulatory regions. Unlike protein-coding variation, rare non-coding genetic variation was almost as likely to increase as decrease protein levels. The regions we identified overlapped predicted tissue-specific enhancers more than promoters, suggesting they represent tissue-specific regulatory regions. Our results have important implications for the identification, and role, of rare non-coding variation associated with common human phenotypes.

genetics↗

Identification and analysis of individuals who deviate from their genetically-predicted phenotype

Findings from genome-wide association studies have facilitated the generation of genetic predictors for many common human phenotypes. Stratifying individuals misaligned to a genetic predictor based on common variants may be important for follow-up studies that aim to identify alternative causal factors. Using genome-wide imputed genetic data, we aimed to classify 158,951 unrelated individuals from the UK Biobank as either concordant or deviating from two well-measured phenotypes. We first applied our methods to standing height: our primary analysis classified 244 individuals (0.15%) as misaligned to their genetically predicted height. We show that these individuals are enriched for self-reporting being shorter or taller than average at age 10, diagnosed congenital malformations, and rare loss-of-function variants in genes previously catalogued as causal for growth disorders. Secondly, we apply our methods to LDL cholesterol. We classified 156 (0.12%) individuals as misaligned to their genetically predicted LDL cholesterol and show that these individuals were enriched for both clinically actionable cardiovascular risk factors and rare genetic variants in genes previously shown to be involved in metabolic processes. Individuals whose LDL-C was higher than expected based on the genetic predictor were also at higher risk of developing coronary artery disease and type-two diabetes, even after adjustment for measured LDL-C, BMI and age, suggesting upward deviation from genetically predicted LDL-C is indicative of generally poor health. Our results remained broadly consistent when performing sensitivity analysis based on a variety of parametric and non-parametric methods to define individuals deviating from polygenic expectation. Our analyses demonstrate the potential importance of quantitatively identifying individuals for further follow-up based on deviation from genetic predictions. Author SummaryHuman genetics is becoming increasingly useful to help predict human traits across a population owing to findings from large-scale genetic association studies and advances in the power of genetic predictors. This provides an opportunity to potentially identify individuals that deviate from genetic predictions for a common phenotype under investigation. For example, an individual may be genetically predicted to be tall, but be shorter than expected. It is potentially important to identify individuals who deviate from genetic predictions as this can facilitate further follow-up to assess likely causes. Using 158,951 unrelated individuals from the UK Biobank, with height and LDL cholesterol, as exemplar traits, we demonstrate that approximately 0.15% & 0.12% of individuals deviate from their genetically predicted phenotypes respectively. We observed these individuals to be enriched for a range of rare clinical diagnoses, as well as rare genetic factors that may be causal. Our analyses also demonstrate several methods for detecting individuals who deviate from genetic predictions that can be applied to a range of continuous human phenotypes.

genetics↗