Search bioRxivSearch

Biology subjects

Stahl, E. A.

Publications and source records attributed to Stahl, E. A..

10 recordsLinked to original sources

Prioritizing risk genes for neurodevelopmental disorders using pathway information

Trio family and case-control studies of next-generation sequencing data have proven integral to understanding the contribution of rare inherited and de novo single-nucleotide variants to the genetic architecture of complex disease. Ideally, such studies should identify individual risk genes of moderate to large effect size to generate novel treatment hypotheses for further follow-up. However, due to insufficient power, gene set enrichment analyses have come to be relied upon for detecting differences between cases and controls, implicating sets of hundreds of genes rather than specific targets for further investigation. Here, we present a Bayesian statistical framework, termed gTADA, that integrates gene-set membership information with gene-level de novo and rare inherited case-control counts, to prioritize risk genes with excess rare variant burden within enriched gene sets. Applying gTADA to available whole-exome sequencing datasets for several neuropsychiatric conditions, we replicated previously reported gene set enrichments and identified novel risk genes. For epilepsy, gTADA prioritized 40 risk genes (posterior probabilities > 0.95), 6 of which replicate in an independent whole-genome sequencing study. In addition, 30/40 genes are novel genes. We found that epilepsy genes had high protein-protein interaction (PPI) network connectivity, and show specific expression during human brain development. Some of the top prioritized EPI genes were connected to a PPI subnetwork of immune genes and show specific expression in prenatal microglia. We also identified multiple enriched drug-target gene sets for EPI which included immunostimulants as well as known antiepileptics. Immune biology was supported specifically by case-control variants from familial epilepsies rather than do novo mutations in generalized encephalitic epilepsy.

genomics

Contribution of rare copy number variants to bipolar disorder risk is limited to schizoaffective cases

BackgroundGenetic risk for bipolar disorder (BD) is conferred through many common alleles, while a role for rare copy number variants (CNVs) is less clear. BD subtypes schizoaffective disorder bipolar type (SAB), bipolar I disorder (BD I) and bipolar II disorder (BD II) differ according to the prominence and timing of psychosis, mania and depression. The factors contributing to the combination of symptoms within a given patient are poorly understood.\n\nMethodsRare, large CNVs were analyzed in 6353 BD cases (3833 BD I [2676 with psychosis, 850 without psychosis], 1436 BD II, 579 SAB) and 8656 controls. Measures of CNV burden were integrated with polygenic risk scores (PRS) for schizophrenia (SCZ) to evaluate the relative contributions of rare and common variants to psychosis risk.\n\nResultsCNV burden did not differ in BD relative to controls when treated as a single diagnostic entity. Burden in SAB was increased compared to controls (p-value = 0.001), BD I (p-value = 0.0003) and BD II (p-value = 0.0007). Burden and SCZ PRS were higher in SAB compared to BD I with psychosis (CNV p-value = 0.0007, PRS p-value = 0.004) and BD I without psychosis (CNV p-value = 0.0004, PRS p-value = 3.9 x 10-5). Within BD I, psychosis was associated with higher SCZ PRS (p-value = 0.005) but not with CNV burden.\n\nConclusionsCNV burden in BD is limited to SAB. Rare and common genetic variants may contribute differently to risk for psychosis and perhaps other classes of psychiatric symptoms.

genetics

mTADA: a framework for analyzing de novo mutations in multiple traits

Joint analysis of multiple traits can result in the identification of associations not found through the analysis of each trait in isolation. In addition, approaches that consider multiple traits can aid in the characterization of shared genetic etiology among those traits. In recent years, parent-offspring trio studies have reported an enrichment of de novo mutations (DNMs) in neuropsychiatric disorders. The analysis of DNM data in the context of neuropsychiatric disorders has implicated multiple putatively causal genes, and a number of reported genes are shared across disorders. However, a joint analysis method designed to integrate de novo mutation data from multiple studies has yet to be implemented. We here introduce multi pi e-trait TAD A (mTADA) which jointly analyzes two traits using DNMs from non-overlapping family samples. mTADA uses two single-trait analysis data sets to estimate the proportion of overlapping risk genes, and reports genes shared between and specific to the relevant disorders. We applied mTADA to >13,000 trios for six disorders: schizophrenia (SCZ), autism spectrum disorder (ASD), developmental disorders (DD), intellectual disability (ID), epilepsy (EPI), and congenital heart disease (CHD). We report the proportion of overlapping risk genes and the specific risk genes shared for each pair of disorders. A total of 153 genes were found to be shared in at least one pair of disorders. The largest percentages of shared risk genes were observed for pairs of DD, ID, ASD, and CHD (>20%) whereas SCZ, CHD, and EPI did not show strong overlaps In risk gene set between them. Furthermore, mTADA identified additional SCZ, EPI and CHD risk genes through integration with DD de novo mutation data. For CHD, using DD information, 31 risk genes with posterior probabilities > 0.8 were identified, and 20 of these 31 genes were not in the list of known CHD genes. We find evidence that most significant CHD risk genes are strongly expressed in prenatal stages of the human genes. Finally, we validated our findings for CHD and EPI in independent cohorts comprising 1241 CHD trios, 226 CHD singletons and 197 EPI trios. Multiple novel risk genes identified by mTADA also had de novo mutations in these independent data sets. The joint analysis method introduced here, mTADA, is able to identify risk genes shared by two traits as well as additional risk genes not found through single-trait analysis only. A number of risk genes reported by mTADA are identified only through joint analysis, specifically when ASD, DD, or ID are one of the two traits examined. This suggests that novel genes for the trait or a new trait might converge to a core gene list of the three traits.

genomics

Identifying tissues implicated in Anorexia Nervosa using Transcriptomic Imputation

Anorexia nervosa (AN) is a complex and serious eating disorder, occurring in ~1% of individuals. Despite having the highest mortality rate of any psychiatric disorder, little is known about the aetiology of AN, and few effective treatments exist.\n\nGlobal efforts to collect large sample sizes of individuals with AN have been highly successful, and a recent study consequently identified the first genome-wide significant locus involved in AN. This result, coupled with other recent studies and epidemiological evidence, suggest that previous characterizations of AN as a purely psychiatric disorder are over-simplified. Rather, both neurological and metabolic pathways may also be involved.\n\nIn order to elucidate more of the system-specific aetiology of AN, we applied transcriptomic imputation methods to 3,495 cases and 10,982 controls, collected by the Eating Disorders Working Group of the Psychiatric Genomics Consortium (PGC-ED). Transcriptomic Imputation (TI) methods approaches use machine-learning methods to impute tissue-specific gene expression from large genotype data using curated eQTL reference panels. These offer an exciting opportunity to compare gene associations across neurological and metabolic tissues. Here, we applied CommonMind Consortium (CMC) and GTEx-derived gene expression prediction models for 13 brain tissues and 12 tissues with potential metabolic involvement (adipose, adrenal gland, 2 colon, 3 esophagus, liver, pancreas, small intestine, spleen, stomach).\n\nWe identified 35 significant gene-tissue associations within the large chromosome 12 region described in the recent PGC-ED GWAS. We applied forward stepwise conditional analyses and FINEMAP to associations within this locus to identify putatively causal signals. We identified four independently associated genes; RPS26, C12orf49, SUOX, and RDH16. We also identified two further genome-wide significant gene-tissue associations, both in brain tissues; REEP5, in the dorso-lateral pre-frontal cortex (DLPFC; p=8.52x10-07), and CUL3, in the caudate basal ganglia (p=1.8x10-06). These genes are significantly enriched for associations with anthropometric phenotypes in the UK BioBank, as well as multiple psychiatric, addiction, and appetite/satiety pathways. Our results support a model of AN risk influenced by both metabolic and psychiatric factors.

genetics

Genome-wide association study implicates CHRNA2 in cannabis use disorder

Introductory paragraphCannabis is the most frequently used illicit psychoactive substance worldwide1. Life time use has been reported among 35-40% of adults in Denmark2 and the United States3. Cannabis use is increasing in the population4-6 and among users around 9% become dependent7. The genetic risk component is high with heritability estimates of 518-70%9. Here we report the first genome-wide significant risk locus for cannabis use disorder (CUD, P=9.31x10-12) that replicates in an independent population (Preplication=3.27x10-3, Pmetaanalysis=9.09x10-12). The finding is based on a genome-wide association study (GWAS) of 2,387 cases and 48,985 controls followed by replication in 5,501 cases and 301,041 controls. The index SNP (rs56372821) is a strong eQTL for CHRNA2 and analyses of the genetic regulated gene expressions identified significant association of CHRNA2 expression in cerebellum with CUD. This indicates a potential therapeutic use in CUD of compounds with agonistic effect on the neuronal acetylcholine receptor alpha-2 subunit encoded by CHRNA2. At the polygenic level analyses revealed a significant decrease in the risk of CUD with increased load of variants associated with cognitive performance.

genomics

Gene expression imputation across multiple brain regions reveals schizophrenia risk throughout development.

Transcriptomic imputation approaches offer an opportunity to test associations between disease and gene expression in otherwise inaccessible tissues, such as brain, by combining eQTL reference panels with large-scale genotype data. These genic associations could elucidate signals in complex GWAS loci and may disentangle the role of different tissues in disease development. Here, we use the largest eQTL reference panel for the dorso-lateral pre-frontal cortex (DLPFC), collected by the CommonMind Consortium, to create a set of gene expression predictors and demonstrate their utility. We applied these predictors to 40,299 schizophrenia cases and 65,264 matched controls, constituting the largest transcriptomic imputation study of schizophrenia to date. We also computed predicted gene expression levels for 12 additional brain regions, using publicly available predictor models from GTEx. We identified 413 genic associations across 13 brain regions. Stepwise conditioning across the genes and tissues identified 71 associated genes (67 outside the MHC), with the majority of associations found in the DLPFC, and of which 14/67 genes did not fall within previously genome-wide significant loci. We identified 36 significantly enriched pathways, including hexosaminidase-A deficiency, and multiple pathways associated with porphyric disorders. We investigated developmental expression patterns for all 67 non-MHC associated genes using BRAINSPAN, and identified groups of genes expressed specifically pre-natally or post-natally.

genetics

Transcriptomic Imputation of Bipolar Disorder and Bipolar subtypes reveals 29 novel associated genes

Bipolar disorder is a complex neuropsychiatric disorder presenting with episodic mood disturbances. In this study we use a transcriptomic imputation approach to identify novel genes and pathways associated with bipolar disorder, as well as three diagnostically and genetically distinct subtypes. Transcriptomic imputation approaches leverage well-curated and publicly available eQTL reference panels to create gene-expression prediction models, which may then be applied to \"impute\" genetically regulated gene expression (GREX) in large GWAS datasets. By testing for association between phenotype and GREX, rather than genotype, we hope to identify more biologically interpretable associations, and thus elucidate more of the genetic architecture of bipolar disorder.\n\nWe applied GREX prediction models for 13 brain regions (derived from CommonMind Consortium and GTEx eQTL reference panels) to 21,488 bipolar cases and 54,303 matched controls, constituting the largest transcriptomic imputation study of bipolar disorder (BPD) to date. Additionally, we analyzed three specific BPD subtypes, including 14,938 individuals with subtype 1 (BD-I), 3,543 individuals with subtype 2 (BD-II), and 1,500 individuals with schizoaffective subtype (SAB).\n\nWe identified 125 gene-tissue associations with BPD, of which 53 represent independent associations after FINEMAP analysis. 29/53 associations were novel; i.e., did not lie within 1Mb of a locus identified in the recent PGC-BD GWAS. We identified 37 independent BD-I gene-tissue associations (10 novel), 2 BD-II associations, and 2 SAB associations. Our BPD, BD-I and BD-II associations were significantly more likely to be differentially expressed in post-mortem brain tissue of BPD, BD-I and BD-II cases than we might expect by chance. Together with our pathway analysis, our results support long-standing hypotheses about bipolar disorder risk, including a role for oxidative stress and mitochondrial dysfunction, the post-synaptic density, and an enrichment of circadian rhythm and clock genes within our results.

genetics

Transcriptional signatures of schizophrenia in hiPSC-derived NPCs and neurons are concordant with signatures from post mortem adult brains

Whereas highly penetrant variants have proven well-suited to human induced pluripotent stem cell (hiPSC)-based models, the power of hiPSC-based studies to resolve the much smaller effects of common variants within the size of cohorts that can be realistically assembled remains uncertain. In developing a large case/control schizophrenia (SZ) hiPSC-derived cohort of neural progenitor cells and neurons, we identified and accounted for a variety of technical and biological sources of variation. Reducing the stochastic effects of the differentiation process by correcting for cell type composition boosted the SZ signal in hiPSC-based models and increased the concordance with post mortem datasets. Because this concordance was strongest in hiPSC-neurons, it suggests that this cell type may better model genetic risk for SZ. We predict a growing convergence between hiPSC and post mortem studies as both approaches expand to larger cohort sizes. For studies of complex genetic disorders, to maximize the power of hiPSC cohorts currently feasible, in most cases and whenever possible, we recommend expanding the number of individuals even at the expense of the number of replicate hiPSC clones.

genetics

Bayesian Integrated Analysis Of Multiple Types Of Rare Variants To Infer Risk Genes For Schizophrenia And Other Neurodevelopmental Disorders

BackgroundIntegrating rare variation from trio family and case/control studies has successfully implicated specific genes contributing to risk of neurodevelopmental disorders (NDDs) including autism spectrum disorders (ASD), intellectual disability (ID), developmental disorders (DD), and epilepsy (EPI). For schizophrenia (SCZ), however, while sets of genes have been implicated through study of rare variation, only two risk genes have been identified.\n\nMethodsWe used hierarchical Bayesian modeling of rare variant genetic architecture to estimate mean effect sizes and risk-gene proportions, analyzing the largest available collection of whole exome sequence (WES) data for schizophrenia (1,077 trios, 6,699 cases and 13,028 controls), and data for four NDDs (ASD, ID, DD, and EPI; total 10,792 trios, and 4,058 cases and controls).\n\nResultsFor SCZ, we estimate 1,551 risk genes, more risk genes and weaker effects than for NDDs. We provide power analyses to predict the number of risk gene discoveries as more data become available, demonstrating greater value of case-control over trio samples. We confirm and augment prior risk gene and gene set enrichment results for SCZ and NDDs. In particular, we detected 98 new DD risk genes at FDR < 0.05. Correlations of risk-gene posterior probabilities are high across four NDDs ({rho} > 0.55), but low between SCZ and the NDDs ({rho} < 0.3). In depth analysis of 288 NDD genes shows highly significant protein-protein interaction (PPI) network connectivity, and functionally distinct PPI subnetworks based on pathway enrichments, single-cell RNA-seq (scRNAseq) cell types and multi-region developmental brain RNA-seq.\n\nConclusionsWe have extended a pipeline used in ASD studies and applied it to infer rare genetic parameters for SCZ and four NDDs. We find many new DD risk genes, supported by gene set enrichment and PPI network connectivity analyses. We find greater similarity among NDDs than between NDDs and SCZ. NDD gene subnetworks are implicated in postnatally expressed presynaptic and postsynaptic genes, and for transcriptional and post-transcriptional gene regulation in prenatal neural progenitor and stem cells.

genomics

Co-localization of Conditional eQTL and GWAS Signatures in Schizophrenia

Causal genes and variants within genome-wide association study (GWAS) loci can be identified by integrating GWAS statistics with expression quantitative trait loci (eQTL) and determining which SNPs underlie both GWAS and eQTL signals. Most analyses, however, consider only the marginal eQTL signal, rather than dissecting this signal into multiple independent eQTL for each gene. Here we show that analyzing conditional eQTL signatures, which could be important under specific cellular or temporal contexts, leads to improved fine mapping of GWAS associations. Using genotypes and gene expression levels from post-mortem human brain samples (N=467) reported by the CommonMind Consortium (CMC), we find that conditional eQTL are widespread; 63% of genes with primary eQTL also have conditional eQTL. In addition, genomic features associated with conditional eQTL are consistent with context specific (i.e. tissue, cell type, or developmental time point specific) regulation of gene expression. Integrating the Psychiatric Genomics Consortium schizophrenia (SCZ) GWAS and CMC conditional eQTL data reveals forty loci with strong evidence for co-localization (posterior probability >0.8), including six loci with co-localization of conditional eQTL. Our co-localization analyses support previously reported genes and identify novel genes for schizophrenia risk, and provide specific hypotheses for their functional follow-up.

genetics