Search bioRxiv⌕ Search

Biology subjects

Horsfield, S. T.

Publications and source records attributed to Horsfield, S. T..

2 recordsLinked to original sources

Graph-based Nanopore Adaptive Sampling with GNASTy enables sensitive pneumococcal serotyping in complex samples

Serotype surveillance of Streptococcus pneumoniae (the pneumococcus) is critical for understanding the effectiveness of current vaccination strategies. However, existing methods for serotyping are limited in their ability to identify co-carriage of multiple pneumococci and detect novel serotypes. To develop a scalable and portable serotyping method that overcomes these challenges, we employed Nanopore Adaptive Sampling (NAS), an on-sequencer enrichment method which selects for target DNA in real-time, for direct detection of S. pneumoniae in complex samples. Whereas NAS targeting the whole S. pneumoniae genome was ineffective in the presence of non-pathogenic streptococci, the method was both specific and sensitive when targeting the capsular biosynthetic locus (CBL), the operon that determines S. pneumoniae serotype. NAS significantly improved coverage and yield of the CBL relative to sequencing without NAS, and accurately quantified the relative prevalence of serotypes in samples representing co-carriage. To maximise the sensitivity of NAS to detect novel serotypes, we developed and benchmarked a new pangenome-graph algorithm, named GNASTy. We show that GNASTy outperforms the current NAS implementation, which is based on linear genome alignment, when a sample contains a serotype absent from the database of targeted sequences. The methods developed in this work provide an improved approach for novel serotype discovery and routine S. pneumoniae surveillance that is fast, accurate and feasible in low resource settings. GNASTy therefore has the potential to increase the density and coverage of global pneumococcal surveillance. One sentence summaryPangenome graph-based Nanopore Adaptive Sampling, presented in our tool GNASTy, is a sensitive, portable and cost-effective method for Streptococcus pneumoniae surveillance.

genomics↗

Accurate and fast graph-based pangenome annotation and clustering with ggCaller

Bacterial genomes differ in both gene content and sequence mutations, which can cause important clinical phenotypic differences such as vaccine escape or antimicrobial resistance. To identify and quantify important variants, all genes within a population must be predicted, functionally annotated and clustered, representing the pangenome. Despite the volume of genome data available, gene prediction and annotation are currently conducted in isolation on individual genomes, which is computationally inefficient and frequently inconsistent across genomes. Here, we introduce the open-source software graph-gene-caller (ggCaller; https://github.com/samhorsfield96/ggCaller). ggCaller combines gene prediction, functional annotation and clustering into a single step using population-wide de Bruijn Graphs, removing redundancy in gene annotation, and resulting in more accurate gene predictions and orthologue clustering. We applied ggCaller to simulated and real-world bacterial genome datasets, comparing it to current state-of-the-art tools. ggCaller is ~50x faster with equivalent or greater accuracy, particularly in datasets with complex sources of error, such as assembly contamination or fragmentation. ggCaller is also an important extension to bacterial genome-wide association studies, enabling querying of annotated graphs for functional analyses. We highlight this application by functionally annotating DNA sequences with significant associations to tetracycline and macrolide resistance in Streptococcus pneumoniae, identifying key resistance determinants that were missed when using only a single reference genome. ggCaller is a novel bacterial genome analysis tool with applications in bacterial epidemiology and evolutionary study.

bioinformatics↗