Search bioRxivSearch

Biology subjects

Davis, J. J.

Publications and source records attributed to Davis, J. J..

4 recordsLinked to original sources

Using machine learning to predict antimicrobial minimum inhibitory concentrations and associated genomic features for nontyphoidal Salmonella

Nontyphoidal Salmonella species are the leading bacterial cause of food-borne disease in the United States. Whole genome sequences and paired antimicrobial susceptibility data are available for Salmonella strains because of surveillance efforts from public health agencies. In this study, a collection of 5,278 nontyphoidal Salmonella genomes, collected over 15 years in the United States, were used to generate XGBoost-based machine learning models for predicting minimum inhibitory concentrations (MICs) for 15 antibiotics. The MIC prediction models have average accuracies between 95-96% within {+/-} 1 two-fold dilution factor and can predict MICs with no a priori information about the underlying gene content or resistance phenotypes of the strains. By selecting diverse genomes for training sets, we show that highly accurate MIC prediction models can be generated with fewer than 500 genomes. We also show that our approach for predicting MICs is stable over time despite annual fluctuations in antimicrobial resistance gene content in the sampled genomes. Finally, using feature selection, we explore the important genomic regions identified by the models for predicting MICs. To date, this is one of the largest MIC modeling studies to be published. Our strategy for developing whole genome sequence-based models for surveillance and clinical diagnostics can be readily applied to other important human pathogens.

bioinformatics

Panaconda: Application of pan-synteny graph models to genome content analysis

MotivationWhole-genome alignment and pan-genome analysis are useful tools in understanding the similarities and differences of many genomes in an evolutionary context. Here we introduce the concept of pan-synteny graphs, an analysis method that combines elements of both to represent conservation and change of multiple prokaryotic genomes at an architectural level. Pan-synteny graphs represent a reference free approach for the comparison of many genomes and allows for the identification of synteny, insertion, deletion, replacement, inversion, recombination, missed assembly joins, evolutionary hotspots, and reference based scaffolding.\n\nResultsWe present an algorithm for creating whole genome multiple sequence comparisons and a model for representing the similarities and differences among sequences as a graph of syntenic gene families. As part of the pan-synteny graph creation, we first create a de Bruijn graph. Instead of the alphabet of nucleotides commonly used in genome assembly, we use an alphabet of gene families. This de Bruijn graph is then processed to create the pan-synteny graph. Our approach is novel in that it explicitly controls how regions from the same sequence and genome are aligned and generates a graph in which all sequences are fully represented as paths. This method harnesses previous computation involved in protein family calculation to speed up the creation of whole genome alignment for many genomes. We provide the software suite Panaconda, for the calculation of pan-synteny graphs given annotation input, and an implementation of methods for their layout and visualization.\n\nAvailabilityPanaconda is available at https://github.com/aswarren/pangenome_graphs and datasets used in examples are available at https://github.com/aswarren/pangenome_examples\n\nContactAndrew Warren anwarren@vt.edu

bioinformatics

Discovery and whole genome sequencing of a human clinical isolate of the novel species Klebsiella quasivariicola sp. nov.

Originally thought to be a single species, Klebsiella pneumoniae has been divided into three distinct species: K. pneumoniae, K. quasipneumoniae and K. variicola. In a recent study of 1,777 extended-spectrum beta-lactamase (ESBL)-producing Klebsiella strains recovered from human infections in Houston, we discovered one strain (KPN1705) causing a wound infection that was phylogenetically distinct from all currently recognized Klebsiella species. Whole genome sequencing of strain KPN1705 revealed that it was single locus variant of the multilocus sequence type ST-1155. This sequence type was reported only once previously. To further investigate the phylogeny of these two organisms, we sequenced the genome of strain KPN1705 to closure and compared its genetic features to Klebsiella reference strains. Results demonstrated strain KPN1705 extensively shares core gene content, antimicrobial resistance genes, and plasmids with K. pneumoniae, K. quasipneumoniae and K. variicola. Since strain KPN1705 and the previously reported novel strain are phylogenetically most closely related to K. variicola, we propose the name K. quasivariicola sp. nov.

microbiology

The DOE Systems Biology Knowledgebase (KBase)

The U.S. Department of Energy Systems Biology Knowledgebase (KBase) is an open-source software and data platform designed to meet the grand challenge of systems biology -- predicting and designing biological function from the biomolecular (small scale) to the ecological (large scale). KBase is available for anyone to use, and enables researchers to collaboratively generate, test, compare, and share hypotheses about biological functions; perform large-scale analyses on scalable computing infrastructure; and combine experimental evidence and conclusions that lead to accurate models of plant and microbial physiology and community dynamics. The KBase platform has (1) extensible analytical capabilities that currently include genome assembly, annotation, ontology assignment, comparative genomics, transcriptomics, and metabolic modeling; (2) a web-browser-based user interface that supports building, sharing, and publishing reproducible and well-annotated analyses with integrated data; (3) access to extensive computational resources; and (4) a software development kit allowing the community to add functionality to the system.

bioinformatics