Search bioRxiv⌕ Search

Biology subjects

Somineni, H.

Publications and source records attributed to Somineni, H..

4 recordsLinked to original sources

Causal considerations can determine the utility of machine learning assisted GWAS

Machine Learning (ML) is increasingly employed to generate phenotypes for genetic discovery, either by imputing existing phenotypes into larger cohorts or by creating novel phenotypes. While these ML-derived phenotypes can significantly increase sample size, and thereby empower genetic discovery, they can also inflate the false discovery rate (FDR). Recent research has focused on developing estimators that leverage both true and machine-learned phenotypes to properly control the type-I error. Our work complements these efforts by exploring how the true positive rate (TPR) and FDR depend on the causal relationships among the inputs to the ML model, the true phenotypes, and the environment. Using a simulation-based framework, we study architectures in which the machine-learned proxy phenotype is derived from biomarkers (i.e. inputs) either causally upstream or downstream of the target phenotype. We show that no inflation of the false discovery rate occurs when the proxy phenotype is generated from upstream biomarkers, but that false discoveries can occur when the proxy phenotype is generated from downstream biomarkers. Next, we show that power to detect variants truly associated with the target phenotype depends on its heritability and correlation with the proxy phenotype. However, the source of the correlation is key to evaluating a proxy phenotype's utility for genetic discovery. We demonstrate that evaluating machine-learned proxy phenotypes using out-of-sample predictive performance (e.g. phenotypic correlation) provides a poor lens on utility. This is because overall predictive performance does not differentiate between genetic and environmental correlation. In addition to parsing these properties of machine-learned phenotypes via simulations, we further illustrate them using real-world data from the UK Biobank.

genetics↗

EmbedGEM: A framework to evaluate the utility of embeddings for genetic discovery

Machine learning (ML)-derived embeddings are a compressed representation of high content data modalities. Embeddings can capture detailed information about disease states and have been qualitatively shown to be useful in genetic discovery. Despite their promise, embeddings have a major limitation: it is unclear if genetic variants associated with embeddings are relevant to the disease or trait of interest. In this work we describe EmbedGEM (Embedding Genetic Evaluation Methods), a framework to systematically evaluate the utility of embeddings in genetic discovery. EmbedGEM focuses on comparing embeddings along two axes: heritability and disease relevance. As measures of heritability, we consider the number of genome-wide significant associations and the mean{chi} 2 statistic at significant loci. For disease relevance, we compute polygenic risk scores for each embedding principal component, then evaluate their association with high-confidence disease or trait labels in a held-out evaluation patient set. While our development of EmbedGEM is motivated by embeddings, the approach is generally applicable to multivariate traits, and can readily be extended to accommodate additional metrics along the evaluation axes. We demonstrate EmbedGEMs utility by evaluating embeddings and multivariate traits in two separate datasets: i) a synthetic dataset simulated to demonstrate the ability of the framework to correctly rank traits based on their heritability and disease relevance, and ii) a real data from the UK Biobank including metabolic and liver-related traits. Importantly, we show that greater disease relevance does not automatically follow from greater heritability.

bioinformatics↗

Pitfalls in performing genome-wide association studies on ratio traits

Genome-wide association studies (GWAS) are often performed on ratios composed of a numerator trait divided by a denominator trait. Examples include body mass index (BMI) and the waist-to-hip ratio, among many others. Explicitly or implicitly, the goal of forming the ratio is typically to adjust for an association between the numerator and denominator. While forming ratios may be clinically expedient, there are several important issues with performing GWAS on ratios. Forming a ratio does not "adjust" for the denominator in the sense of conditioning on it, and it is unclear whether associations with ratios are attributable to the numerator, the denominator, or both. Here we demonstrate that associations arising in ratio GWAS can be entirely denominator-driven, implying that at least some associations uncovered by ratio GWAS may be due solely to a putative adjustment variable. In a survey of 10 common ratio traits, we find that the ratio model disagrees with the adjusted model (performing GWAS on the numerator while conditioning on the denominator) at around 1/3 of loci. Using BMI as an example, we show that variants detected by only the ratio model are more strongly associated with the denominator (height), while variants detected by only the adjusted model are more strongly associated with the numerator (weight). Although the adjusted model provides effect sizes with a clearer interpretation, it is susceptible to collider bias. We propose and validate a simple method of correcting for the genetic component of collider bias via leave-one-chromosome-out polygenic scoring.

genetics↗

An allelic series rare variant association test for candidate gene discovery

Allelic series are of candidate therapeutic interest due to the existence of a dose-response relationship between the functionality of a gene and the degree or severity of a phenotype. We define an allelic series as a gene in which increasingly deleterious mutations lead to increasingly large phenotypic effects, and develop a gene-based rare variant association test specifically targeted for the identification of allelic series. Building on the well-known burden and sequence kernel association (SKAT) tests, we specify a variety of association models, covering different genetic architectures, and integrate these into a COding-variant Allelic Series Test (COAST). Through extensive simulations, we confirm that COAST maintains the type I error and improves power when the pattern of coding-variant effect sizes increases monotonically with mutational severity. We applied COAST to identify allelic series for 4 circulating lipid traits and 5 cell count traits among 145,735 subjects with available whole exome sequencing data from the UK Biobank. Compared with optimal SKAT (SKAT-O), COAST identified 29% more Bonferroni significant associations with circulating lipid traits, on average, and 82% more with cell count traits. All of the gene-trait associations identified by COAST have corroborating evidence either from rare-variant associations in the full cohort (Genebass, N = 400K), or from common variant associations in the GWAS catalog. In addition to detecting many gene-trait associations present in Genebass using only a fraction (36.9%) of the sample, COAST detects associations, such as ANGPTL4 with triglycerides, that are absent from Genebass but which have clear common variant support.

genetics↗