Search bioRxivSearch

Biology subjects

Mudge, J. M.

Publications and source records attributed to Mudge, J. M..

3 recordsLinked to original sources

Systematic re-annotation of 191 genes associated with early-onset epilepsy unmasks de novo variants linked to Dravet syndrome in novel SCN1A exons

The early infantile epileptic encephalopathies (EIEE) are a group of rare, severe neurodevelopmental disorders, where even the most thorough sequencing studies leave 60-65% of patients without a molecular diagnosis. Here, we explore the incompleteness of transcript models used for exome and genome analysis as one potential explanation for lack of current diagnoses. Therefore, we have updated the GENCODE gene annotation for 191 epilepsy-associated genes, using human brain-derived transcriptomic libraries and other data to build 3,550 novel putative transcript models. The extended transcriptional footprint of these genes allowed for 294 intronic or intergenic variants, found in human mutation databases, to be reclassified as exonic, while a further 70 intronic variants were reclassified as splice-site proximal. Using SCN1A as a case study due to its close phenotype/genotype correlation with Dravet syndrome, we screened 122 people with Dravet syndrome, or a similar phenotype, with a panel of novel exon sequences representing eight established genes and identified two de novo SCN1A variants that now, through improved gene annotation can be ascribed to residing among novel exons. These two (from 122 screened patients, 1.6%) new molecular diagnoses carry significant clinical implications. Furthermore, we identified a previously-classified SCN1A intronic Dravet-associated variant that now lies within a deeply conserved novel exon. Our findings illustrate the potential gains of thorough gene annotation in improving diagnostic yields for genetic disorders. We would expect to find new molecular diagnoses in our 191 genes that were originally suspected by clinicians for patients, with a negative diagnosis.

genomics

Transcript expression-aware annotation improves rare variant discovery and interpretation

The acceleration of DNA sequencing in patients and population samples has resulted in unprecedented catalogues of human genetic variation, but the interpretation of rare genetic variants discovered using such technologies remains extremely challenging. A striking example of this challenge is the existence of disruptive variants in dosage-sensitive disease genes, even in apparently healthy individuals. Through manual curation of putative loss of function (pLoF) variants in haploinsufficient disease genes in the Genome Aggregation Database (gnomAD)(1), we show that one explanation for this paradox involves alternative mRNA splicing, which allows exons of a gene to be expressed at varying levels across cell types. Currently, no existing annotation tool systematically incorporates this exon expression information into variant interpretation. Here, we develop a transcript-level annotation metric, the proportion expressed across transcripts (pext), which summarizes isoform quantifications for variants. We calculate this metric using 11,706 tissue samples from the Genotype Tissue Expression project(2) (GTEx) and show that it clearly differentiates between weakly and highly evolutionarily conserved exons, a proxy for functional importance. We demonstrate that expression-based annotation selectively filters 22.8% of falsely annotated pLoF variants found in haploinsufficient disease genes in gnomAD, while removing less than 4% of high-confidence pathogenic variants in the same genes. Finally, we apply our expression filter to the analysis of de novo variants in patients with autism spectrum disorder (ASD) and developmental disorders and intellectual disability (DD/ID) to show that pLoF variants in weakly expressed regions have effect sizes similar to those of synonymous variants, while pLoF variants in highly expressed exons are most strongly enriched among cases versus controls. Our annotation is fast, flexible, and generalizable, making it possible for any variant file to be annotated with any isoform expression dataset, and will be valuable for rare disease diagnosis, rare variant burden analyses in complex disorders, and curation and prioritization of variants in recall-by-genotype studies.

genomics

Molecular complexity of the major urinary protein system of the Norway rat, Rattus norvegicus

Major urinary proteins (MUP) are the major component of the urinary protein fraction in house mice (Mus spp.) and rats (Rattus spp.). The structure, polymorphism and functions of these lipocalins have been well described in the western European house mouse (Mus musculus domesticus), clarifying their role in semiochemical communication. The complexity of these roles in the mouse raises the question of similar functions in other rodents, including the Norway rat, Rattus norvegicus. Norway rats express MUPs in urine but information about specific MUP isoform sequences and functions is limited. In this study, we present a detailed molecular characterization of the MUP proteoforms expressed in the urine of two laboratory strains, Wistar Han and Brown Norway, and wild caught animals, using a combination of manual gene annotation, intact protein mass spectrometry and bottom-up mass spectrometry-based proteomic approaches. Detailed sequencing of the proteins reveals a less complex pattern of primary sequence polymorphism than the mouse. However, unlike the mouse, rat MUPs exhibit added complexity in the form of post-translational modifications including phosphorylation and exoproteolytic trimming of specific isoforms. The possibility that urinary MUPs may have different roles in rat chemical communication than those they play in the house mouse is also discussed.

biochemistry