Search bioRxivSearch

Biology subjects

Buxbaum, J.

Publications and source records attributed to Buxbaum, J..

5 recordsLinked to original sources

Prioritizing risk genes for neurodevelopmental disorders using pathway information

Trio family and case-control studies of next-generation sequencing data have proven integral to understanding the contribution of rare inherited and de novo single-nucleotide variants to the genetic architecture of complex disease. Ideally, such studies should identify individual risk genes of moderate to large effect size to generate novel treatment hypotheses for further follow-up. However, due to insufficient power, gene set enrichment analyses have come to be relied upon for detecting differences between cases and controls, implicating sets of hundreds of genes rather than specific targets for further investigation. Here, we present a Bayesian statistical framework, termed gTADA, that integrates gene-set membership information with gene-level de novo and rare inherited case-control counts, to prioritize risk genes with excess rare variant burden within enriched gene sets. Applying gTADA to available whole-exome sequencing datasets for several neuropsychiatric conditions, we replicated previously reported gene set enrichments and identified novel risk genes. For epilepsy, gTADA prioritized 40 risk genes (posterior probabilities > 0.95), 6 of which replicate in an independent whole-genome sequencing study. In addition, 30/40 genes are novel genes. We found that epilepsy genes had high protein-protein interaction (PPI) network connectivity, and show specific expression during human brain development. Some of the top prioritized EPI genes were connected to a PPI subnetwork of immune genes and show specific expression in prenatal microglia. We also identified multiple enriched drug-target gene sets for EPI which included immunostimulants as well as known antiepileptics. Immune biology was supported specifically by case-control variants from familial epilepsies rather than do novo mutations in generalized encephalitic epilepsy.

genomics

mTADA: a framework for analyzing de novo mutations in multiple traits

Joint analysis of multiple traits can result in the identification of associations not found through the analysis of each trait in isolation. In addition, approaches that consider multiple traits can aid in the characterization of shared genetic etiology among those traits. In recent years, parent-offspring trio studies have reported an enrichment of de novo mutations (DNMs) in neuropsychiatric disorders. The analysis of DNM data in the context of neuropsychiatric disorders has implicated multiple putatively causal genes, and a number of reported genes are shared across disorders. However, a joint analysis method designed to integrate de novo mutation data from multiple studies has yet to be implemented. We here introduce multi pi e-trait TAD A (mTADA) which jointly analyzes two traits using DNMs from non-overlapping family samples. mTADA uses two single-trait analysis data sets to estimate the proportion of overlapping risk genes, and reports genes shared between and specific to the relevant disorders. We applied mTADA to >13,000 trios for six disorders: schizophrenia (SCZ), autism spectrum disorder (ASD), developmental disorders (DD), intellectual disability (ID), epilepsy (EPI), and congenital heart disease (CHD). We report the proportion of overlapping risk genes and the specific risk genes shared for each pair of disorders. A total of 153 genes were found to be shared in at least one pair of disorders. The largest percentages of shared risk genes were observed for pairs of DD, ID, ASD, and CHD (>20%) whereas SCZ, CHD, and EPI did not show strong overlaps In risk gene set between them. Furthermore, mTADA identified additional SCZ, EPI and CHD risk genes through integration with DD de novo mutation data. For CHD, using DD information, 31 risk genes with posterior probabilities > 0.8 were identified, and 20 of these 31 genes were not in the list of known CHD genes. We find evidence that most significant CHD risk genes are strongly expressed in prenatal stages of the human genes. Finally, we validated our findings for CHD and EPI in independent cohorts comprising 1241 CHD trios, 226 CHD singletons and 197 EPI trios. Multiple novel risk genes identified by mTADA also had de novo mutations in these independent data sets. The joint analysis method introduced here, mTADA, is able to identify risk genes shared by two traits as well as additional risk genes not found through single-trait analysis only. A number of risk genes reported by mTADA are identified only through joint analysis, specifically when ASD, DD, or ID are one of the two traits examined. This suggests that novel genes for the trait or a new trait might converge to a core gene list of the three traits.

genomics

Elevated polygenic burden for autism is associated with differential DNA methylation at birth.

BackgroundAutism spectrum disorder (ASD) is a severe neurodevelopmental disorder characterized by deficits in social communication and restricted, repetitive behaviors, interests, or activities. The etiology of ASD involves both inherited and environmental risk factors, with epigenetic processes hypothesized as one mechanism by which both genetic and non-genetic variation influence gene regulation and pathogenesis.\n\nMethodsWe quantified neonatal methylomic variation in 1,263 infants - of whom ~50% went on to subsequently develop ASD - using DNA isolated from a unique collection of archived blood spots taken shortly after birth. We used matched genetic data from the same individuals to examine the molecular consequences of ASD genetic risk variants, identifying methylomic variation associated with elevated polygenic burden for ASD. In addition, we performed DNA methylation quantitative trait loci (mQTL) mapping to prioritize target genes from ASD GWAS findings.\n\nResultsAlthough we did not identify specific loci showing consistent changes in neonatal DNA methylation associated with later ASD, we found a significant association between increased polygenic burden for autism and methylomic variation at two CpG sites located proximal to a robust GWAS signal for ASD on chromosome 8.\n\nConclusionsThis study is the largest analysis of DNA methylation in ASD yet undertaken and the first to integrate both genetic and epigenetic variation at birth in ASD. We demonstrate the utility of using a polygenic risk score to identify molecular variation associated with disease, and of using mQTL to refine the functional and regulatory variation associated with ASD risk variants.

genetics

Bayesian Integrated Analysis Of Multiple Types Of Rare Variants To Infer Risk Genes For Schizophrenia And Other Neurodevelopmental Disorders

BackgroundIntegrating rare variation from trio family and case/control studies has successfully implicated specific genes contributing to risk of neurodevelopmental disorders (NDDs) including autism spectrum disorders (ASD), intellectual disability (ID), developmental disorders (DD), and epilepsy (EPI). For schizophrenia (SCZ), however, while sets of genes have been implicated through study of rare variation, only two risk genes have been identified.\n\nMethodsWe used hierarchical Bayesian modeling of rare variant genetic architecture to estimate mean effect sizes and risk-gene proportions, analyzing the largest available collection of whole exome sequence (WES) data for schizophrenia (1,077 trios, 6,699 cases and 13,028 controls), and data for four NDDs (ASD, ID, DD, and EPI; total 10,792 trios, and 4,058 cases and controls).\n\nResultsFor SCZ, we estimate 1,551 risk genes, more risk genes and weaker effects than for NDDs. We provide power analyses to predict the number of risk gene discoveries as more data become available, demonstrating greater value of case-control over trio samples. We confirm and augment prior risk gene and gene set enrichment results for SCZ and NDDs. In particular, we detected 98 new DD risk genes at FDR < 0.05. Correlations of risk-gene posterior probabilities are high across four NDDs ({rho} > 0.55), but low between SCZ and the NDDs ({rho} < 0.3). In depth analysis of 288 NDD genes shows highly significant protein-protein interaction (PPI) network connectivity, and functionally distinct PPI subnetworks based on pathway enrichments, single-cell RNA-seq (scRNAseq) cell types and multi-region developmental brain RNA-seq.\n\nConclusionsWe have extended a pipeline used in ASD studies and applied it to infer rare genetic parameters for SCZ and four NDDs. We find many new DD risk genes, supported by gene set enrichment and PPI network connectivity analyses. We find greater similarity among NDDs than between NDDs and SCZ. NDD gene subnetworks are implicated in postnatally expressed presynaptic and postsynaptic genes, and for transcriptional and post-transcriptional gene regulation in prenatal neural progenitor and stem cells.

genomics

New mutations, old statistical challenges

Based on targeted sequencing of 208 genes in 11,730 neurodevelopmental disorder cases, Stessman et al. report the identification of 91 genes associated (at a False Discovery Rate [FDR] of 0.1) with autism spectrum disorders (ASD), intellectual disability (ID), and developmental delay (DD)--including what they characterize as 38 novel genes, not previously reported as connected with these diseases1.\n\nIf true, this would represent a substantial step forward. Unfortunately, each of the two discovery analyses (1. De novo mutation analysis and, 2. a comparison of private mutations with public control data) contain critical statistical flaws. When one accounts for these problems, fewer than half of the genes--and very few, if any, of the novel findings--survive. These errors have implications for how future analyses should be conducted, for understanding the genetic basis of these disorders, and for genomic medicine.\n\nWe discuss the two main ana ...

genetics