Search bioRxivSearch

Biology subjects

Kathryn E Holt

Publications and source records attributed to Kathryn E Holt.

9 recordsLinked to original sources

Identification of Klebsiella capsule synthesis loci from whole genome data

BackgroundKlebsiella pneumoniae and close relatives are a growing cause of healthcare-associated infections for which increasing rates of multi-drug resistance are a major concern. The Klebsiella polysaccharide capsule is a major virulence determinant and epidemiological marker. However, little is known about capsule epidemiology since serological typing is not widely accessible, and many isolates are serologically non-typeable. Molecular methods for capsular typing are needed, but existing methods lack sensitivity and specificity and fail to take advantage of the information available in whole-genome sequence data, which is increasingly being generated for surveillance and investigation of Klebsiella.\n\nMethodsWe investigated the diversity of capsule synthesis loci (K loci) among a large, diverse collection of 2503 genome sequences of K. pneumoniae and closely related species. We incorporated analyses of both full-length K locus DNA sequences and clustered protein coding sequences to identify, annotate and compare K locus structures, and we propose a novel method for identifying K loci based on full locus information extracted from whole genome sequences.\n\nResultsA total of 134 distinct K loci were identified, including 31 novel types. Comparative analysis of K locus gene content detected 508 unique protein coding gene clusters that appear to reassort via homologous recombination, generating novel K locus types. Extensive nucleotide diversity was detected among the wzi and wzc genes, both within and between K loci, indicating that current typing schemes based on these genes are inadequate. As a solution, we introduce Kaptive, a novel software tool that automates the process of identifying K loci from large sets of Klebsiella genomes based on full locus information.\n\nConclusionsThis work highlights the extensive diversity of Klebsiella K loci and the proteins that they encode. We propose a standardised K locus nomenclature for Klebsiella, present a curated reference database of all known K loci, and introduce a tool for identifying K loci from genome data (https://github.com/katholt/Kaptive). These developments constitute important new resources for the Klebsiella community for use in genomic surveillance and epidemiology.

Genomics

Genome-scale rates of evolutionary change in bacteria

Estimating the rates at which bacterial genomes evolve is critical to understanding major evolutionary and ecological processes such as disease emergence, long-term host-pathogen associations, and short-term transmission patterns. The surge in bacterial genomic data sets provides a new opportunity to estimate these rates and reveal the factors that shape bacterial evolutionary dynamics. For many organisms estimates of evolutionary rate display an inverse association with the time-scale over which the data are sampled. However, this relationship remains unexplored in bacteria due to the difficulty in estimating genome-wide evolutionary rates, which are impacted by the extent of temporal structure in the data and the prevalence of recombination. We collected 36 whole genome sequence data sets from 16 species of bacterial pathogens to systematically estimate and compare their evolutionary rates and assess the extent of temporal structure in the absence of recombination. The majority (28/36) of data sets possessed sufficient clock-like structure to robustly estimate evolutionary rates. However, in some species reliable estimates were not possible even with \"ancient DNA\" data sampled over many centuries, suggesting that they evolve very slowly or that they display extensive rate variation among lineages. The robustly estimated evolutionary rates spanned several orders of magnitude, from 10-6 to 10-8 nucleotide substitutions site-1 year-1. This variation was largely attributable to sampling time, which was strongly negatively associated with estimated evolutionary rates, with this relationship best described by an exponential decay curve. To avoid potential estimation biases such time-dependency should be considered when inferring evolutionary time-scales in bacteria.

Evolutionary Biology

South Asia as a reservoir for the global spread of ciprofloxacin resistant Shigella sonnei

BackgroundAntimicrobial resistance is a major issue in the Shigellae, particularly as a specific multidrug resistant (MDR) lineage of Shigella sonnei (lineage III) is becoming globally dominant. Ciprofloxacin is a recommended treatment for Shigella infections. However, ciprofloxacin resistant S. sonnei are being increasingly isolated in Asia, and sporadically reported on other continents.\n\nMethods and FindingsHypothesising that Asia is the hub for the recent international spread of ciprofloxacin resistant S. sonnei, we performed whole genome sequencing on a collection of contemporaneous ciprofloxacin resistant S. sonnei isolated in six countries from within and outside of Asia. We reconstructed the recent evolutionary history of these organisms and combined these data with their geographical location of isolation. Placing these sequences into a global phylogeny we found that all ciprofloxacin resistant S. sonnei formed a single clade within a Central Asian expansion of Lineage III. Further, our data show that resistance to ciprofloxacin within S. sonnei can be globally attributed to a single clonal emergence event, encompassing sequential gyrA-S83L, parC-S80I and gyrA-D87G mutations. Geographical data predict that South Asia is the likely primary source of these organisms, which are being regularly exported across Asia and intercontinentally into Australia, the USA and Europe.\n\nConclusionsThis study shows that a single clone, which is widespread in South Asia, is driving the current intercontinental surge of ciprofloxacin resistant S. sonnei and is capable of establishing endemic transmission in new locations. Despite being limited in geographical scope, our work has major implications for understanding the international transfer of antimicrobial resistant S. sonnei, and provides a tractable model for studying how antimicrobial resistant Gram-negative community acquired pathogens spread globally.

Microbiology

Inducible colistin resistance via a disrupted plasmid-borne mcr-1 gene in a 2008 Vietnamese Shigella sonnei isolate

The mcr-1 gene, which confers resistance against the last-resort antimicrobial colistin, was recently discovered in Enterobacteriaceae circulating in China. Through genome sequencing we identified a plasmid-associated inactive form of mcr-1 in a 2008 Vietnamese isolate of Shigella sonnei. The plasmid was conjugated into E. coli and mcr-1 was activated upon exposure to colistin, suggesting the gene has been circulating in human-restricted pathogens for some time but carries a selective fitness cost.

Microbiology

Bandage: interactive visualisation of de novo genome assemblies

SummaryWhile de novo assembly graphs contain assembled contigs (nodes), the connections between those contigs (edges) are difficult for users to access. Bandage (a Bioinformatics Application for Navigating De novo Assembly Graphs Easily) is a tool for visualising assembly graphs with connections. Users can zoom in to specific areas of the graph and interact with it by moving nodes, adding labels, changing colours and extracting sequences. BLAST searches can be performed within the Bandage GUI and the hits are displayed as highlights in the graph. By displaying connections between contigs, Bandage presents new possibilities for analysing de novo assemblies that are not possible through investigation of contigs alone.\n\nAvailability and implementationSource code and binaries are freely available at https://github.com/rrwick/Bandage. Bandage is implemented in C++ and supported on Linux, OS X and Windows.\n\nContactrrwick@gmail.com\n\nSupplementary informationA full feature list and screenshots are available at Bioinformatics online and http://rrwick.github.io/Bandage.

Bioinformatics

ISMapper: Identifying insertion sequences in bacterial genomes from short read sequence data

BackgroundInsertion sequences (IS) are small transposable elements, commonly found in bacterial genomes. Identifying the location of IS in bacterial genomes can be useful for a variety of purposes including epidemiological tracking and predicting antibiotic resistance. However IS are commonly present in multiple copies in a single genome, which complicates genome assembly and the identification of IS insertion sites. Here we present ISMapper, a mapping-based tool for identification of the site and orientation of IS insertions in bacterial genomes, direct from paired-end short read data.\n\nResultsISMapper was validated using three types of short read data: (i) simulated reads from a variety of species, (ii) Illumina reads from 5 isolates for which finished genome sequences were available for comparison, and (iii) Illumina reads from 7 Acinetobacter baumannii isolates for which predicted IS locations were tested using PCR. A total of 20 genomes, including 13 species and 32 distinct IS, were used for validation. ISMapper correctly identified 96% of known IS insertions in the analysis of simulated reads, and 98% in real Illumina reads. Subsampling of real Illumina reads to lower depths indicated ISMapper was reliable for average genome-wide read depths >20x. All ISAba1 insertions identified by ISMapper in the A. baumannii genomes were confirmed by PCR. In each A. baumannii genome, ISMapper successfully identified an IS insertion upstream of the ampC beta-lactamase that could explain phenotypic resistance to third-generation cephalosporins. The utility of ISMapper was further demonstrated by profiling genome-wide IS6110 insertions in 138 publicly available Mycobacterium tuberculosis genomes, revealing lineage-speific inserction and multi inserction hotspot.\n\nConclusionsISMapper provides a rapid and robust method for identifying IS insertion sites direct from short read data, with a high degree of accuracy demonstrated across a wide range of bacteria.

Bioinformatics

The infant airway microbiome in health and disease impacts later asthma development

The nasopharynx (NP) is a reservoir for microbes associated with acute respiratory illnesses (ARI). The development of asthma is initiated during infancy, driven by airway inflammation associated with infections. Here, we report viral and bacterial community profiling of NP aspirates across a birth cohort, capturing all lower respiratory illnesses during their first year. Most infants were initially colonized with Staphylococcus or Corynebacterium before stable colonization with Alloiococcus or Moraxella, with transient incursions of Streptococcus, Moraxella or Haemophilus marking virus-associated ARIs. Our data identify the NP microbiome as a determinant for infection spread to the lower airways, severity of accompanying inflammatory symptoms, and risk for future asthma development. Early asymptomatic colonization with Streptococcus was a strong asthma predictor, and antibiotic usage disrupted asymptomatic colonization patterns.

Genomics

Extensive capsule locus variation and large-scale genomic recombination within the Klebsiella pneumoniae clonal complex 258/11.

Klebsiella pneumoniae clonal complex (CC) 258/11, comprising sequence types (STs) 258, 11 and closely related STs, is associated with dissemination of the K. pneumoniae carbapenemase (KPC). Hospital outbreaks of KPC CC258/11 infections have been observed globally and are very difficult to treat. As a consequence there is renewed interest in alternative infection control measures such as vaccines and phage or depolymerase treatments targeting the K pneumoniae polysaccharide capsule. To date, 78 immunologically distinct capsule variants have been described in K. pneumoniae. Previous investigations of ST258 and a small number of closely related strains suggested capsular variation was limited within this clone; only two distinct ST258 capsular synthesis (cps) loci have been identified, both acquired through large-scale recombination events (>50 kbp). Here we report comparative genomic analysis of the broader K. pneumoniae CC258/11. Our data indicate that several large-scale recombination events have shaped the genomes of CC258/11, and that definition of the complex should be broadened to include ST395 (also reported to harbour KPC). We identified 11 different cps loci within CC258/11, suggesting that capsular switching is actually common within the complex. We also observed several insertion sequences (IS) within the cps loci, and show further diversification of two loci through IS activity. These findings suggest the capsular loci of clinically important K. pneumoniae are under diversifying selection, which alters our understanding of the evolution of this important clone and has implications for the design of control measures targeting the capsule.

Genomics

SRST2: Rapid genomic surveillance for public health and hospital microbiology labs

Rapid molecular typing of bacterial pathogens is critical for public health epidemiology, surveillance and infection control, yet routine use of whole genome sequencing (WGS) for these purposes poses significant challenges. Here we present SRST2, a read mapping-based tool for fast and accurate detection of genes, alleles and multi-locus sequence types (MLST) from WGS data. Using >900 genomes from common pathogens, we show SRST2 is highly accurate and outperforms assembly-based methods in terms of both gene detection and allele assignment. Here we have demonstrated the use of SRST2 for microbial genome surveillance in a variety of public health and hospital settings. In the face of rising threats of antimicrobial resistance and emerging virulence amongst bacterial pathogens, SRST2 represents a powerful tool for rapidly extracting clinically useful information from raw WGS data. Source code is available from http://katholt.github.io/srst2/.

Genomics