Search bioRxivSearch

Biology subjects

Kelly, B. J.

Publications and source records attributed to Kelly, B. J..

3 recordsLinked to original sources

Global Analysis of Human mRNA Folding Disruptions in Synonymous Variants Demonstrates Significant Population Constraint

BackgroundIn most organisms the structure of an mRNA molecule is crucial in determining speed of translation, half-life, splicing propensities and final protein configuration. Synonymous variants which distort this wildtype mRNA structure may be pathogenic as a consequence. However, current clinical guidelines classify synonymous or "silent" single nucleotide variants (sSNVs) as largely benign unless a role in RNA splicing can be demonstrated. ResultsWe developed novel software to conduct a global transcriptome study in which RNA folding statistics were computed for 469 million SNVs in 45,800 transcripts using an Apache Spark implementation of ViennaRNA in the cloud. Focusing our analysis on the subset of 17.9 million sSNVs, we discover that variants predicted to disrupt mRNA structure have lower rates of incidence in the human population. Given that the community lacks tools to evaluate the potential pathogenic impact of sSNVs, we introduce a "Structural Predictivity Index" (SPI) to quantify this constraint due to mRNA structure. ConclusionsOur findings support the hypothesis that sSNVs may play a role in genetic disorders due to their effects on mRNA structure. Our RNA-folding scores provide a means of gauging the structural constraint operating on any sSNV in the human genome. Given that the majority of patients with rare or as yet to be diagnosed disease lack a molecular diagnosis, these scores have the potential to enable discovery of novel genetic etiologies. Our RNA Stability Pipeline as well as ViennaRNA structural metrics and SPI scores for all human synonymous variants can be downloaded from GitHub https://github.com/nch-igm/rna-stability.

genomics

Gut microbiome features associated with Clostridium difficile colonization in puppies

In people, colonization with Clostridium difficile, the leading cause of antibiotic-associated diarrhea, has been shown to be associated with distinct gut microbial features, including reduced bacterial community diversity and depletion of key taxa. In dogs, the gut microbiome features that define C. difficile colonization are less well understood. We sought to define the gut microbiome features associated with C. difficile colonization in puppies, a population where the prevalence of C. difficile has been shown to be elevated, and to define the effect of puppy age and litter upon these features and C. difficile risk. We collected fecal samples from weaned (n=27) and unweaned (n=74) puppies from 13 litters and analyzed the effects of colonization status, age and litter on microbial diversity using linear mixed effects models.\n\nColonization with C. difficile was significantly associated with younger age, and colonized puppies had significantly decreased bacterial community diversity and differentially abundant taxa compared to non-colonized puppies, even when adjusting for age. C. difficile colonization remained associated with decreased bacterial community diversity, but the association did not reach statistical significance in a mixed effects model incorporating litter as a random effect.\n\nEven though litter explained a greater proportion (67%) of the variability in microbial diversity than colonization status, we nevertheless observed heterogeneity in gut microbial community diversity and colonization status within more than half of the litters, suggesting that the gut microbiome contributes to colonization resistance against C. difficile. The colonization of puppies with C. difficile has important implications for the potential zoonotic transfer of this organism to people. The identified associations point to mechanisms by which C. difficile colonization may be reduced.

epidemiology

Samovar: Single-sample mosaic SNV calling with linked reads

We present Samovar, a mosaic single-nucleotide variant (SNV) caller for linked-read whole-genome shotgun sequencing data. Samovar scores candidate sites using a random forest model trained using the input dataset that considers read quality, phasing, and linked-read characteristics. We show Samovar calls mosaic SNVs within a single sample with accuracy comparable to what previously required trios or matched tumor/normal pairs and outperform single-sample mosaic variant callers at MAF 5%-50% with at least 30x coverage. Furthermore, we use Samovar to find somatic variants in whole genome sequencing of both tumor and normal from 13 pediatric cancer cases that can be corroborated with high recall with whole exome sequencing. Samovar is available open-source at https://github.com/cdarby/samovar under the MIT license.

genomics