Search bioRxivSearch

Biology subjects

Hsiao, W.

Publications and source records attributed to Hsiao, W..

3 recordsLinked to original sources

MentaLiST - A fast MLST caller for large wgMLST schemes

MLST (multi-locus sequence typing) is a classic technique for genotyping bacteria, widely applied for pathogen outbreak surveillance. Traditionally, MLST is based on identifying sequence types from a small number of housekeeping genes. With the increasing availability of whole-genome sequencing (WGS) data, MLST methods have evolved toward larger typing schemes, based on a few hundred genes (core genome MLST, cgMLST) to a few thousand genes (whole genome MLST, wgMLST). Such large-scale MLST schemes have been shown to provide a finer resolution and are increasingly used in various contexts such as hospital outbreaks or foodborne pathogen outbreaks. This methodological shift raises new computational challenges, especially given the large size of the schemes involved. Very few available MLST callers are currently capable of dealing with large MLST schemes.\n\nWe introduce MentaLiST, a new MLST caller, based on a k-mer voting algorithm and written in the Julia language, specifically designed and implemented to handle large typing schemes. We test it on real and simulated data to show that MentaLiST is faster than any other available MLST caller while providing the same or better accuracy, and is capable of dealing with MLST scheme with up to thousands of genes while requiring limited computational resources. MentaLiST source code and easy installation instructions using a Conda package are available at https://github.com/WGS-TB/MentaLiST.

bioinformatics

Antibiotic resistance genes in agriculture and urban influenced watersheds in southwestern British Columbia

BackgroundThe dissemination of antibiotic resistance genes (ARGs) from anthropogenic activities into the environment poses an emerging public health threat. Water constitutes a major vehicle for transport of both biological material and chemical substances. The present study focused on putative antibiotic resistance and integrase genes present in the microbiome of agricultural, urban influenced and protected watersheds in southwestern British Columbia, Canada. A metagenomics approach and high throughput quantitative PCR (HT qPCR) were used to screen for elements of resistance including ARGs and integron-associated integrase genes (intI). Sequencing of bacterial genomic DNA was used to characterize the resistome of microbial communities present in watersheds over a one-year period.\n\nResultsData mining using CARD and Integrall databases enabled the identification of putative antibiotic resistance genes present in watershed samples. Antibiotic resistance genes presence in samples from various watershed locations was low relative to the microbial population (<1 %). Analysis of the metagenomic sequences detected a total of 78 ARGs and intI1 across all watershed locations. The relative abundance and richness of antibiotic resistance genes was found to be highest in agriculture impacted watersheds compared to protected and urban watersheds. Gene copy numbers (GCNs) from a subset of 21 different elements of antibiotic resistance were further estimated using HT qPCR. Most GCNs of ARGs were found to be variable over time. A downstream transport pattern was observed in the impacted watersheds (urban and agricultural) during dry months. Urban and agriculture impacted sites had a higher GCNs of ARGs compared to protected sites. Similar to other reports, this study found a strong association between intI1 and ARGs (e.g., sul1), an association which may be used as a proxy for anthropogenic activities. Chemical analysis of water samples for three major groups of antibiotics was negative. However, the high richness and GCNs of ARGs in impacted sites suggest effects of effluents on microbial communities are occurring even at low concentrations of antimicrobials in the water column.\n\nConclusionAntibiotic resistance and integrase genes in a year-long metagenomic study showed that ARGs were driven mainly by environmental factors from anthropogenized sites in agriculture and urban watersheds. Environmental factors accounted for almost 40% of the variability observed in watershed locations.

microbiology

SNVPhyl: A Single Nucleotide Variant Phylogenomics pipeline for microbial genomic epidemiology

MotivationThe recent widespread application of whole-genome sequencing (WGS) for microbial disease investigations has spurred the development of new bioinformatics tools, including a notable proliferation of phylogenomics pipelines designed for infectious disease surveillance and outbreak investigation. Transitioning the use of WGS data out of the research lab and into the front lines of surveillance and outbreak response requires user-friendly, reproducible, and scalable pipelines that have been well validated.\n\nResultsSNVPhyl (Single Nucleotide Variant Phylogenomics) is a bioinformatics pipeline for identifying high-quality SNVs and constructing a whole genome phylogeny from a collection of WGS reads and a reference genome. Individual pipeline components are integrated into the Galaxy bioinformatics framework, enabling data analysis in a user-friendly, reproducible, and scalable environment. We show that SNVPhyl can detect SNVs with high sensitivity and specificity and identify and remove regions of high SNV density (indicative of recombination). SNVPhyl is able to correctly distinguish outbreak from non-outbreak isolates across a range of variant-calling settings, sequencing-coverage thresholds, or in the presence of contamination.\n\nAvailabilitySNVPhyl is available as a Galaxy workflow, Docker and virtual machine images, and a Unix-based command-line application. SNVPhyl is released under the Apache 2.0 license and available at http://snvphyl.readthedocs.io/ or at https://github.com/phac-nml/snvphyl-galaxy.

bioinformatics