Search bioRxivSearch

Biology subjects

Dahl, A.

Publications and source records attributed to Dahl, A..

10 recordsLinked to original sources

Reverse GWAS: Using Genetics to Identify and Model Phenotypic Subtypes

Recent and classical work has revealed biologically and medically significant subtypes in complex diseases and traits. However, relevant subtypes are often unknown, unmeasured, or actively debated, making automatic statistical approaches to subtype definition particularly valuable. We propose reverse GWAS (RGWAS) to identify and validate subtypes using genetics and multiple traits: while GWAS seeks the genetic basis of a given trait, RGWAS seeks to define trait subtypes with distinct genetic bases. Unlike existing approaches relying on off-the-shelf clustering methods, RGWAS uses a bespoke decomposition, MFMR, to model covariates, binary traits, and population structure. We use extensive simulations to show these features can be crucial for power and calibration. We validate RGWAS in practice by recovering known stress subtypes in major depressive disorder. We then show the utility of RGWAS by identifying three novel subtypes of metabolic traits. We biologically validate these metabolic subtypes with SNP-level tests and a novel polygenic test: the former recover known metabolic GxE SNPs; the latter suggests genetic heterogeneity may explain substantial missing heritability. Crucially, statins, which are widely prescribed and theorized to increase diabetes risk, have opposing effects on blood glucose across metabolic subtypes, suggesting potential have potential translational value.\n\nAuthor summaryComplex diseases depend on interactions between many known and unknown genetic and environmental factors. However, most studies aggregate these strata and test for associations on average across samples, though biological factors and medical interventions can have dramatically different effects on different people. Further, more-sophisticated models are often infeasible because relevant sources of heterogeneity are not generally known a priori. We introduce Reverse GWAS to simultaneously split samples into homogeneoues subtypes and to learn differences in genetic or treatment effects between subtypes. Unlike existing approaches to computational subtype identification using high-dimensional trait data, RGWAS accounts for covariates, binary disease traits and, especially, population structure; these features are each invaluable in extensive simulations. We validate RGWAS by recovering known genetic subtypes of major depression. We demonstrate RGWAS is practically useful in a metabolic study, finding three novel subtypes with both SNP- and polygenic-level heterogeneity. Importantly, RGWAS can uncover differential treatment response: for example, we show that statin, a common drug and potential type 2 diabetes risk factor, may have opposing subtype-specific effects on blood glucose.

genetics

Existence and implications of population variance structure

Identifying the genetic and environmental factors underlying phenotypic differences between populations is fundamental to multiple research communities. To date, studies have focused on the relationship between population and phenotypic mean. Here we consider the relationship between population and phenotypic variance, i.e., \"population variance structure.\" In addition to gene-gene and gene-environment interaction, we show that population variance structure is a direct consequence of natural selection. We develop the ancestry double generalized linear model (ADGLM), a statistical framework to jointly model population mean and variance effects. We apply ADGLM to several deeply phenotyped datasets and observe ancestry-variance associations with 12 of 44 tested traits in ~113K British individuals and 3 of 14 tested traits in ~3K Mexican, Puerto Rican, and African-American individuals. We show through extensive simulations that population variance structure can both bias and reduce the power of genetic association studies, even when principal components or linear mixed models are used. ADGLM corrects this bias and improves power relative to previous methods in both simulated and real datasets. Additionally, ADGLM identifies 17 novel genotype-variance associations across six phenotypes.

genetics

GxEMM: Extending linear mixed models to general gene-environment interactions

Gene-environment interaction (GxE) is a well-known source of non-additive inheritance. GxE can be important in applications ranging from basic functional genomics to precision medical treatment. Further, GxE effects elude inherently-linear LMMs and may explain missing heritability. We propose a simple, unifying mixed model for polygenic interactions (GxEMM) to capture the aggregate effect of small GxE effects spread across the genome. GxEMM extends existing LMMs for GxE in two important ways. First, it extends to arbitrary environmental variables, not just categorical groups. Second, GxEMM can estimate and test for environment-specific heritability. In simulations where the assumptions of existing methods do not hold, we show that GxEMM improves estimates of ordinary and GxE heritability and increases power to test for polygenic GxE. We then use GxEMM to prove that the heritability of major depression (MD) is reduced by stress, which we previously conjectured but could not prove with prior methods, and that a tail of polygenic GxE effects remains unexplained by MD GWAS.

genetics

GBAT: a gene-based association method for robust trans-gene regulation detection

Identification of trans-eQTLs has been limited by a heavy multiple testing burden, read-mapping biases, and hidden confounders. To address these issues, we developed GBAT, a powerful gene-based method that allows robust detection of trans gene regulation. Using simulated and real data, we show that GBAT drastically increases detection of trans-gene regulation over standard trans-eQTL analyses.

genetics

On negative heritability and negative estimates of heritability

AO_SCPLOWBSTRACTC_SCPLOWWe consider the problem of interpreting negative maximum likelihood estimates of heritability that sometimes arise from popular statistical models of additive genetic variation. These may result from random noise acting on estimates of genuinely positive heritability, but we argue that they may also arise from misspecification of the standard additive mechanism that is supposed to justify the statistical procedure. Researchers should be open to the possibility that negative heritability estimates could reflect a real physical feature of the biological process from which the data were sampled.

genetics

Instructive starPEG-Heparin biohybrid 3D cultures for modeling human neural stem cell plasticity, neurogenesis, and neurodegeneration

Three-dimensional models of human neural development and neurodegeneration are crucial when exploring stem-cell-based regenerative therapies in a tissue-mimetic manner. However, existing 3D culture systems are not sufficient to model the inherent plasticity of NSCs due to their ill-defined composition and lack of controllability of the physical properties. Adapting a glycosaminoglycan-based, cell-responsive hydrogel platform, we stimulated primary and induced human neural stem cells (NSCs) to manifest neurogenic plasticity and form extensive neuronal networks in vitro. The 3D cultures exhibited neurotransmitter responsiveness, electrophysiological activity, and tissue-specific extracellular matrix (ECM) deposition. By whole transcriptome sequencing, we identified that 3D cultures express mature neuronal markers, and reflect the in vivo make-up of mature cortical neurons compared to 2D cultures. Thus, our data suggest that our established 3D hydrogel culture supports the tissue-mimetic maturation of human neurons. We also exemplarily modeled neurodegenerative conditions by treating the cultures with A{beta}42 peptide and observed the known human pathological effects of Alzheimers disease including reduced NSC proliferation, impaired neuronal network formation, synaptic loss and failure in ECM deposition as well as elevated Tau hyperphosphorylation and formation of neurofibrillary tangles. We determined the changes in transcriptomes of primary and induced NSC-derived neurons after A{beta}42, providing a useful resource for further studies. Thus, our hydrogel-based human cortical 3D cell culture is a powerful platform for studying various aspects of neural development and neurodegeneration, as exemplified for A{beta}42 toxicity and neurogenic stem cell plasticity.\n\nSignificanceNeural stem cells (NSC) are reservoir for new neurons in human brains, yet they fail to form neurons after neurodegeneration. Therefore, understanding the potential use of NSCs for stem cell-based regenerative therapies requires tissue-mimetic humanized experimental systems. We report the adaptation of a 3D bio-instructive hydrogel culture system where human NSCs form neurons that later form networks in a controlled microenvironment. We also modeled neurodegenerative toxicity by using Amyloid-beta4 peptide, a hallmark of Alzheimers disease, observed phenotypes reminiscent of human brains, and determined the global gene expression changes during development and degeneration of neurons. Thus, our reductionist humanized culture model will be an important tool to address NSC plasticity, neurogenicity, and network formation in health and disease.

neuroscience

Effect of vitamin D supplementation on biomarkers of inflammation and immune function: functional genomics analysis of the BEST-D trial

Vitamin D deficiency has been associated with multiple diseases, but the causal relevance and underlying processes are not fully understood. Elucidating the mechanisms of action of drug treatments in humans is challenging, but application of functional genomic approaches in randomised trials may afford an opportunity to systematically assess molecular responses to treatments. In the Biochemical Efficacy and Safety Trial of Vitamin D (BEST-D), 305 community-dwelling individuals aged over 65 years were randomly allocated to treatment with vitamin D34000 IU, 2000 IU or placebo daily for 12 months. Genome-wide genotypes at baseline, and transcriptome and plasma levels of cytokines (IFN-{gamma}, IL-10, IL-8, IL-6 and TNF-) at baseline and after 12 months, were measured. The trial had >90% power to detect a 2-fold change in gene expression. Allocation to vitamin D for 12-months was associated with 2-fold higher plasma levels of 25-hydroxy-vitamin D (25[OH]D), but had no significant effect on whole-blood gene expression (FDR <5%) or on plasma levels of cytokines compared with placebo. In pre-specified analysis, rs7041 (intron variant, GC) had a significant effect on circulating levels of 25(OH)D in the low dose but not on the placebo or high dose vitamin D regimen. A gene expression quantitative trait locus analysis (eQTL) demonstrated evidence of 31,568 cis-eQTLs (unique SNP-probe pairs) among individuals at baseline and 34,254 after supplementation for 12 months (any dose), but had no significant effect on cis-eQTLs specific to vitamin D supplementation. The trial demonstrates the feasibility of application of functional genomics approaches in randomised trials to assess the effects of vitamin D on immune function.\n\nOne sentence summarySupplementation with high-dose vitamin D in older people for 12 months in a randomised, placebo-controlled trial had no significant effect on gene expression or on plasma concentrations of cytokines.\n\nTrial registrationSRCTN registry (Number 07034656) and the European Clinical Trials Database (EudraCT Number 2011-005763-24).\n\nFundingMedical Research Council, British Heart Foundation, Wellcome Trust, European Research Council and Clinical Trial Service Unit, Nuffield Department of Population Health, University of Oxford, Oxford, United Kingdom\n\nCopyrightOpen access article under the terms of CC BY.

clinical trials

Singleton Variants Dominate the Genetic Architecture of Human Gene Expression

The vast majority of human mutations have minor allele frequencies (MAF) under 1%, with the plurality observed only once (i.e., \"singletons\"). While Mendelian diseases are predominantly caused by rare alleles, their cumulative contribution to complex phenotypes remains largely unknown. We develop and rigorously validate an approach to jointly estimate the contribution of all alleles, including singletons, to phenotypic variation. We apply our approach to transcriptional regulation, an intermediate between genetic variation and complex disease. Using whole genome DNA and lymphoblastoid cell line RNA sequencing data from 360 European individuals, we conservatively estimate that singletons contribute ~25% of cis-heritability across genes (dwarfing the contributions of other frequencies). Strikingly, the majority (~76%) of singleton heritability derives from ultra-rare variants absent from thousands of additional samples. We develop a novel inference procedure to demonstrate that our results are consistent with rampant purifying selection shaping the regulatory architecture of most human genes.

evolutionary biology

Adjusting For Principal Components Of Molecular Phenotypes Induces Replicating False Positives

High-throughput measurements of molecular phenotypes provide an unprecedented opportunity to model cellular processes and their impact on disease. Such highly-structured data is strongly confounded, and principal components and their variants reliably estimate latent confounders. Conditioning on PCs in downstream analyses is known to improve power and reduce multiple-testing miscalibration and is an indispensable element of thousands of published functional genomic analyses. Further clarifying this approach is of fundamental interest to the genomics and statistics communities. We uncover a novel bias induced by PC conditioning and provide an analytic, deterministic and intuitive approximation. The bias exists because PCs are, roughly, unshielded colliders on a causal path: because PCs partially incorporate a causal genotype effect on one phenotype, the genotype becomes correlated with every phenotype conditional on PCs. We empirically quantify this bias in realistic simulations. For small genetic effects, a nearly negligible bias is observed for all tested PC variants. For large genetic effects, or other differential covariates, dramatic false positives can arise. Though one PC variant (supervised SVA) largely avoids this bias, it is computationally prohibitive genome-wide; further, its immunity to this bias is novel. Our analysis informs best practices for confounder correction in genomic studies.

genetics

Statistical properties of simple random-effects models for genetic heritability

Random-effects models are a popular tool for analysing total narrow-sense heritability for simple quantitative phenotypes on the basis of large-scale SNP data. Recently, there have been disputes over the validity of conclusions that may be drawn from such analysis. We derive some of the fundamental statistical properties of heritability estimates arising from these models, showing that the bias will generally be small. We show that that the score function may be manipulated into a form that facilitates intelligible interpretations of the results. We use this score function to explore the behavior of the model when certain key assumptions of the model are not satisfied -- shared environment, measurement error, and genetic effects that are confined to a small subset of sites -- as well as to elucidate the meaning of negative heritability estimates that may arise.\n\nThe variance and bias depend crucially on the variance of certain functionals of the singular values of the genotype matrix. A useful baseline is the singular value distribution associated with genotypes that are completely independent -- that is, with no linkage and no relatedness -- for a given number of individuals and sites. We calculate the corresponding variance and bias for this setting.\n\nMSC 2010 subject classifications: Primary 92D10; secondary 62P10; 62F10; 60B20.

genetics