Search bioRxivSearch

Biology subjects

Niestroj, L.-M.

Publications and source records attributed to Niestroj, L.-M..

2 recordsLinked to original sources

Evaluation of copy number burden in specific epilepsy types from a genome-wide study of 18,564 subjects

Rare and large copy number variants (CNVs) around known genomic hotspots are strongly implicated in epilepsy etiology. But it remains unclear whether the observed associations are specific to an epilepsy phenotype, and if additional risk signal can be found outside hotspots. Here, we present the largest CNV burden and first CNV breakpoint level association analysis in epilepsy to date with 11,246 European epilepsy cases and 7,318 ancestry-matched controls. We studied five epilepsy phenotypes: genetic generalized epilepsy, lesional focal epilepsy, non-acquired focal epilepsy, epileptic encephalopathy, and unclassified epilepsy. We discovered novel epilepsy-associated CNV loci and further characterized the CNV burden enrichment among phenotype-specific epilepsies. Finally, we provide evidence for deletion burden outside of known hotspot regions and show that CNVs play a significant role in the genetic architecture of lesional focal epilepsies.

genetics

Identification of pathogenic variant enriched regions across genes and gene families

Missense variant interpretation is challenging. Essential regions for protein function are conserved among gene family members, and genetic variants within these regions are potentially more likely to confer risk to disease. Here, we generated 2,871 gene family protein sequence alignments involving 9,990 genes and performed missense variant burden analyses to identify novel essential protein regions. We mapped 2,219,811 variants from the general population into these alignments and compared their distribution with 65,034 missense variants from patients. With this gene family approach, we identified 398 regions enriched for patient variants spanning 33,887 amino acids in 1,058 genes. As a comparison, testing the same genes individually we identified less patient variant enriched regions involving only 2,167 amino acids and 180 genes. Next, we selected de novo variants from 6,753 patients with neurodevelopmental disorders and 1,911 unaffected siblings, and observed a 5.56-fold enrichment of patient variants in our identified regions (95% C.I. =2.76-Inf, p-value = 6.66x10-8). Using an independent ClinVar variant set, we found missense variants inside the identified regions are 111-fold more likely to be classified as pathogenic in comparison to benign classification (OR = 111.48, 95% C.I = 68.09-195.58, p-value < 2.2e-16). All patient variant enriched regions identified (PERs) are available online through a user-friendly platform for interactive data mining, visualization and download at http://per.broadinstitute.org. In summary, our gene family burden analysis approach identified novel patient variant enriched regions in protein sequences. This annotation can empower variant interpretation.

genetics