Search bioRxivSearch

Biology subjects

Katz, L. S.

Publications and source records attributed to Katz, L. S..

2 recordsLinked to original sources

Comparison Of Multi-locus Sequence Typing Software For Next Generation Sequencing Data

Multi-locus sequence typing (MLST) is a widely used method for categorising bacteria. Increasingly MLST is being performed using next generation sequencing data by reference labs and for clinical diagnostics. Many software applications have been developed to calculate sequence types from NGS data; however, there has been no comprehensive review to date on these methods. We have compared six of these applications against real and simulated data and present results on: 1. the accuracy of each method against traditional typing methods, 2. the performance on real outbreak datasets, 3. in the impact of contamination and varying depth of coverage, and 4. the computational resource requirements.\n\nDATA SUMMARYO_LISimulated reads for datasets testing coverage and mixed samples have been deposited in Figshare; DOI: https://doi.org/10.6084/m9.figshare.4602301.vl\nC_LIO_LIOutbreak databases are available from Github; url - https://github.com/WGS-standards-and-analysis/datasets\nC_LIO_LIDocker containers used to run each of the applications are available from Github; url - https://tinyurl.com/z7ks2ft\nC_LIO_LIAccession numbers for the data used in this paper are available in the Supplementary material.\nC_LI\n\nWe confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {ballotx}\n\nIMPACT STATEMENTSequence typing is rapidly transitioning from traditional sequencing methods to using whole genome sequencing. A number of in silico prediction methods have been developed on an ad hoc basis and aim to replicate Multi-locus sequence typing (MLST). This is the first study to comprehensively evaluate multiple MLST software applications on real validated datasets and on common simulated difficult cases. It will give researchers a clearer understanding of the accuracy, limitations and computational performance of the methods they use, and will assist future researchers to choose the most appropriate method for their experimental goals.

bioinformatics

SNVPhyl: A Single Nucleotide Variant Phylogenomics pipeline for microbial genomic epidemiology

MotivationThe recent widespread application of whole-genome sequencing (WGS) for microbial disease investigations has spurred the development of new bioinformatics tools, including a notable proliferation of phylogenomics pipelines designed for infectious disease surveillance and outbreak investigation. Transitioning the use of WGS data out of the research lab and into the front lines of surveillance and outbreak response requires user-friendly, reproducible, and scalable pipelines that have been well validated.\n\nResultsSNVPhyl (Single Nucleotide Variant Phylogenomics) is a bioinformatics pipeline for identifying high-quality SNVs and constructing a whole genome phylogeny from a collection of WGS reads and a reference genome. Individual pipeline components are integrated into the Galaxy bioinformatics framework, enabling data analysis in a user-friendly, reproducible, and scalable environment. We show that SNVPhyl can detect SNVs with high sensitivity and specificity and identify and remove regions of high SNV density (indicative of recombination). SNVPhyl is able to correctly distinguish outbreak from non-outbreak isolates across a range of variant-calling settings, sequencing-coverage thresholds, or in the presence of contamination.\n\nAvailabilitySNVPhyl is available as a Galaxy workflow, Docker and virtual machine images, and a Unix-based command-line application. SNVPhyl is released under the Apache 2.0 license and available at http://snvphyl.readthedocs.io/ or at https://github.com/phac-nml/snvphyl-galaxy.

bioinformatics