Search bioRxivSearch

Biology subjects

May, P.

Publications and source records attributed to May, P..

4 recordsLinked to original sources

Isolation of nucleic acids from low biomass samples: detection and removal of sRNA contaminants

Sequencing-based analyses of low-biomass samples are known to be prone to misinterpretation due to the potential presence of contaminating molecules derived from laboratory reagents and environments. Due to its inherent instability, contamination with RNA is usually considered to be unlikely. Here we report the presence of small RNA (sRNA) contaminants in widely used microRNA extraction kits and means for their depletion. Sequencing of sRNAs extracted from human plasma samples was performed and significant levels of non-human (exogenous) sequences were detected. The source of the most abundant of these sequences could be traced to the microRNA extraction columns by qPCR-based analysis of laboratory reagents. The presence of artefactual sequences originating from the confirmed contaminants were furthermore replicated in a range of published datasets. To avoid artefacts in future experiments, several protocols for the removal of the contaminants were elaborated, minimal amounts of starting material for artefact-free analyses were defined, and the reduction of contaminant levels for identification of bona fide sequences using ultra-clean extraction kits was confirmed. In conclusion, this is the first report of the presence of RNA molecules as contaminants in laboratory reagents. The described protocols should be applied in the future to avoid confounding sRNA studies.

molecular biology

Gene family information facilitates variant interpretation and identification of disease-associated genes

Differentiating risk-conferring from benign missense variants, and therefore optimal calculation of gene-variant burden, represent a major challenge in particular for rare and genetic heterogeneous disorders. While orthologous gene conservation is commonly employed in variant annotation, approximately 80% of known disease-associated genes are paralogs and belong to gene families. It has not been thoroughly investigated how gene family information can be utilized for disease gene discovery and variant interpretation. We developed a paralog conservation score to empirically evaluate whether paralog conserved or nonconserved sites of in-human paralogs are important for protein function. Using this score, we demonstrate that disease-associated missense variants are significantly enriched at paralog conserved sites across all disease groups and disease inheritance models tested. Next, we assessed whether gene family information could assist in discovering novel disease-associated genes. We subsequently developed a gene family de novo enrichment framework that identified 43 exome-wide enriched gene families including 98 de novo variant carrying genes in more than 10k neurodevelopmental disorder patients. 33 gene family enriched genes represent novel candidate genes which are brain expressed and variant constrained in neurodevelopmental disorders.

genetics

Reassessment Of Lesion-Associated Gene And Variant Pathogenicity In Focal Human Epilepsies

PurposeIncreasing availability of surgically resected brain tissue from Focal Cortical Dysplasia and low-grade epilepsy-associated tumor patients fostered large-scale genetic examination. However, assessment of germline and somatic variant pathogenicity remains difficult.\n\nMethodsHere, we critically reevaluated the pathogenicity for all neuropathology-associated variants reported to date in the PubMed and ClinVar databases, including 12 disease-related genes and 88 neuropathology-associated missense variants. We (1) assessed evolutionary gene constraint using the pLI and missense z scores, (2) applied guidelines by the American College of Medical Genetics and Genomics (ACMG), and (3) predicted pathogenicity by using PolyPhen-2, CADD, and GERP.\n\nResultsConstraint analysis classified only seven out of 12 genes to be likely disease-associated, while 35 (40%) of those 88 variants were classified as being variants of unknown significance (VUS) and 53 (60%) as being likely pathogenic (LPII). Pathogenicity prediction yielded discrimination between neuropathology-associated variants (LPII and VUS) and rare variant scores obtained from individuals present in the Genome Aggregation Database (gnomAD).\n\nConclusionWe conclude that several VUS are likely disease-associated and will be reclassified by future molecular evidence. In summary, interpretation of lesion-associated gene variants remains complex while the application of current ACMG guidelines including bioinformatic pathogenicity prediction will help improving interpretation and prediction.

genetics

The Spectrum Of De Novo Variants In Neurodevelopmental Disorders With Epilepsy

Epilepsy is a frequent feature of neurodevelopmental disorders (NDD) but little is known about genetic differences between NDD with and without epilepsy. We analyzed de novo variants (DNV) in 6753 parent-offspring trios ascertained for different NDD. In the subset of 1942 individuals with NDD with epilepsy, we identified 33 genes with a significant excess of DNV, of which SNAP25 and GABRB2 had previously only limited evidence for disease association. Joint analysis of all individuals with NDD also implicated CACNA1E as a novel disease gene. Comparing NDD with and without epilepsy, we found missense DNV, DNV in specific genes, age of recruitment and severity of intellectual disability to be associated with epilepsy. We further demonstrate to what extent our results impact current genetic testing as well as treatment, emphasizing the benefit of accurate genetic diagnosis in NDD with epilepsy.

genetics