Search bioRxivSearch

Biology subjects

Bentley, S. D.

Publications and source records attributed to Bentley, S. D..

17 recordsLinked to original sources

Fast Hierarchical Bayesian Analysis of Population Structure

We present fastbaps, a fast solution to the genetic clustering problem. Fastbaps rapidly identifies an approximate fit to a Dirichlet Process Mixture model (DPM) for clustering multilocus genotype data. Our efficient model-based clustering approach is able to cluster datasets 10-100 times larger than the existing model-based methods, which we demonstrate by analysing an alignment of over 110,000 sequences of HIV-1 pol genes. We also provide a method for rapidly partitioning an existing hierarchy in order to maximise the DPM model marginal likelihood, allowing us to split phylogenetic trees into clades and subclades using a population genomic model. Extensive tests on simulated data as well as a diverse set of real bacterial and viral datasets show that fastbaps provides comparable or improved solutions to previous model-based methods, while generally being significantly faster. The method is made freely available under an open source MIT licence as an easy to use R package at https://github.com/gtonkinhill/fastbaps.

genomics

Temporal population structure of invasive Group B Streptococcus during a period of rising disease incidence shows expansion of a CC17 clone

Group B Streptococcus (GBS) is a major cause of neonatal invasive disease worldwide. In the Netherlands, the incidence of the disease increased, despite the introduction of prevention guidelines in 1999. This was accompanied by changes in pathogen genotype distribution, with a significant increase in the prevalence of isolates belonging to clonal complex (CC) 17. To better understand the mechanisms of temporal changes in the epidemiology of GBS genotypes that correlated with the rise in disease incidence, we applied whole genome sequencing (WGS) to study a national collection of invasive GBS isolates. A total of 1345 isolates from patients aged 0 - 89 days and collected between 1987 and 2016 in the Netherlands were sequenced and characterised. The GBS population contained 5 major lineages representing CC17 (39%), CC19 (25%), CC23 (18%), CC10 (9%), and CC1 (7%). There was a significant rise in the prevalence of isolates representing CC17 and CC23 among cases of early-and late-onset disease, due to expansion of discrete sub-lineages. The most prominent was shown by a CC17 sub-lineage, identified here as CC17-1A, which experienced a major clonal expansion at the end of the 1990s. The CC17-1A expansion correlated with the emergence of a novel phage carrying a gene encoding a putative adhesion protein, named here StrP. The first occurrence of this phage (designated phiStag1) within the collection in 1997, was followed by multiple, independent acquisitions by CC17 and parallel clonal expansions of CC17-1A and another cluster, CC17-1B. The CC17-1A clone was identified in external datasets, and represents a globally distributed invasive sub-lineage of CC17. Our work describes how a sudden change in the epidemiology of specific GBS sub-lineages, in particular CC17-1A, correlates with the rise in the disease incidence, and indicates a putative key role of a novel phage in driving the expansion of this CC17 clone.\n\nAuthor summaryGroup B Streptococcus (GBS) is a commensal organism of the gastrointestinal and genitourinary tracts. However, it is also an opportunistic pathogen and a major cause of neonatal invasive disease, which can be classified into early-onset (0 - 6 days of life) or late-onset (7 - 89 days of life). Current disease prevention strategy involves intrapartum antibiotic prophylaxis (IAP), which aims to prevent the transmission of GBS from mother to baby during labour. Many developed countries adapted national IAP guidelines. In the Netherlands, these were introduced in 1999. However, the incidence of GBS disease increased after IAP introduction. In this study we applied whole genome sequencing to characterise a nationwide collection of invasive GBS from cases of neonatal disease that occurred between 1987 and 2016. Analysis of GBS population structure involving phylogenetic partitioning of individual lineages revealed that the rise in disease incidence involved the expansion of specific clusters from two major GBS lineages, CC17 and CC23. Our study provides new insights into the recent evolution of the hypervirulent CC17 and describes a rapid expansion of a discrete, pre-existing sub-lineage that occurred after acquisition of a novel phage carrying a putative adhesion protein gene, underscoring the major role of CC17 in neonatal diseases.

genomics

Prediction of post-vaccine population structure of Streptococcus pneumoniae using accessory gene frequencies

Predicting how pathogen populations will change over time is challenging. Such has been the case with Streptococcus pneumoniae, an important human pathogen, and the pneumococcal conjugate vaccines (PCVs), which target only a fraction of the strains in the population. Here, we use the frequencies of accessory genes to predict changes in the pneumococcal population after vaccination, hypothesizing that these frequencies reflect negative frequency-dependent selection (NFDS) on the gene products. We find that the standardized predicted fitness of a strain estimated by an NFDS-based model at the time the vaccine is introduced enables to predict whether the strain increases or decreases in prevalence following vaccination. Further, we are able to forecast the equilibrium post-vaccine population composition and assess the invasion capacity of emerging lineages. Overall, we provide a method for predicting the impact of an intervention on pneumococcal populations with potential application to other bacterial pathogens in which NFDS is a driving force.

evolutionary biology

Fast and flexible bacterial genomic epidemiology with PopPUNK

The routine use of genomics for disease surveillance provides the opportunity for high-resolution bacterial epidemiology.\n\nHowever, current whole-genome clustering and multi-locus typing approaches do not fully exploit core and accessory genomic variation, and cannot both automatically identify, and subsequently expand, clusters of significantly-similar isolates in large datasets and across species.\n\nHere we describe PopPUNK (Population Partitioning Using Nucleotide K-mers; https://poppunk.readthedocs.io/en/latest/). software implementing scalable and expandable annotation- and alignment-free methods for population analysis and clustering.\n\nVariable-length k-mer comparisons are used to distinguish isolates divergence in shared sequence and gene content, which we demonstrate to be accurate over multiple orders of magnitude using both simulated data and real datasets from ten taxonomically-widespread species. Connections between closely-related isolates of the same strain are robustly identified, despite variation in the discontinuous pairwise distance distributions that reflects species diverse evolutionary patterns. PopPUNK can process 103-104 genomes as single batch, with minimal memory use and runtimes up to 200-fold faster than existing methods. Clusters of strains remain consistent as new batches of genomes are added, which is achieved without needing to re-analyse all genomes de novo.\n\nThis facilitates real-time surveillance with stable cluster naming and allows for outbreak detection using hundreds of genomes in minutes. Interactive visualisation and online publication is streamlined through automatic output of results to multiple platforms.\n\nPopPUNK has been designed as a flexible platform that addresses important issues with currently used whole-genome clustering and typing methods, and has potential uses across bacterial genetics and public health research.

genomics

Bayesian inference of ancestral dates on bacterial phylogenetic trees

The sequencing and comparative analysis of a collection of bacterial genomes from a single species or lineage of interest can lead to key insights into its evolution, ecology or epidemiology. The tool of choice for such a study is often to build a phylogenetic tree, and more specifically when possible a dated phylogeny, in which the dates of all common ancestors are estimated. Here we propose a new Bayesian methodology to construct dated phylogenies which is specifically designed for bacterial genomics. Unlike previous Bayesian methods aimed at building dated phylogenies, we consider that the phylogenetic relationships between the genomes have been previously evaluated using a standard phylogenetic method, which makes our methodology much faster and scalable. This two-steps approach also allows us to directly exploit existing phylogenetic methods that detect bacterial recombination, and therefore to account for the effect of recombination in the construction of a dated phylogeny. We analysed many simulated datasets in order to benchmark the performance of our approach in a wide range of situations. Furthermore, we present applications to three different real datasets from recent bacterial genomic studies. Our methodology is implemented in a R package called BactDating which is freely available for download at https://github.com/xavierdidelot/BactDating.

bioinformatics

Antimicrobial exposure in sexual networks drives divergent evolution in modern gonococci

The sexually transmitted pathogen Neisseria gonorrhoeae is regarded as being on the way to becoming an untreatable superbug. Despite its clinical importance, little is known about its emergence and evolution, and how this corresponds with the introduction of antimicrobials. We present a genome-based phylogeographic analysis of 419 gonococcal isolates from across the globe. Results indicate that modern gonococci originated in Europe or Africa as late as the 16thcentury and subsequently disseminated globally. We provide evidence that the modern gonococcal population has been shaped by antimicrobial treatment of sexually transmitted and other infections, leading to the emergence of two major lineages with different evolutionary strategies. The well-described multi-resistant lineage is associated with high rates of homologous recombination and infection in high-risk sexual networks where antimicrobial treatment is frequent. A second, multi-susceptible lineage associated with heterosexual networks, where asymptomatic infection is more common, was also identified, with potential implications for infection control.

genomics

Global emergence and population dynamics of divergent serotype 3 CC180 pneumococci

Streptococcus pneumoniae serotype 3 remains a significant cause of morbidity and mortality worldwide, despite inclusion in the 13-valent pneumococcal conjugate vaccine (PCV13). Serotype 3 increased in carriage since the implementation of PCV13 in the United States, while invasive disease rates remain unchanged. We investigated the persistence of serotype 3 in carriage and disease, through genomic analyses of a global sample of 301 serotype 3 isolates of the Netherlands3-31 (PMEN31) clone CC180, combined with associated patient data and PCV utilization among countries of isolate collection. We assessed phenotypic variation between dominant clades in capsule charge (zeta potential), capsular polysaccharide shedding, and susceptibility to opsonophagocytic killing, which have previously been associated with carriage duration, invasiveness, and vaccine escape. We identify a recent shift in the CC180 population attributed to a lineage termed Clade II, which was estimated by Bayesian coalescent analysis to have first appeared in 1968 [95% HPD: 1939-1989] and increased in prevalence and effective population size thereafter. Clade II isolates are divergent from the pre-PCV13 serotype 3 population in non-capsular antigenic composition, competence, and antibiotic susceptibility, the last resulting from the acquisition of a Tn916-like conjugative transposon. Differences in recombination rates among clades correlated with variations in the ATP-binding subunit of Clp protease as well as amino acid substitutions in the comCDE operon. Opsonophagocytic killing assays elucidated the low observed efficacy of PCV13 against serotype 3. Variation in PCV13 use among sampled countries was not independently correlated with the CC180 population shift; therefore, genotypic and phenotypic differences in protein antigens and, in particular, antibiotic resistance may have contributed to the increase of Clade II. Our analysis emphasizes the need for routine, representative sampling of isolates from disperse geographic regions, including historically under-sampled areas. We also highlight the value of genomics in resolving antigenic and epidemiological variations within a serotype, which may have implications for future vaccine development.\n\nAuthor SummaryStreptococcus pneumoniae is a leading cause of bacterial pneumoniae, meningitis, and otitis media. Despite inclusion in the most recent pneumococcal conjugate vaccine, PCV13, serotype 3 remains epidemiologically important globally. We investigated the persistence of serotype 3 using whole-genome sequencing data form 301 isolates collected among 24 countries from 1993-2014. Through phylogenetic analysis, we identified three distinct lineages within a single clonal complex, CC180, and found one has recently emerged and grown in prevalence. We then compared genomic difference among lineages as well as variations in pneumococcal vaccine use among sampled countries. We found that the recently emerged lineage, termed Clade II, has a higher prevalence of antibiotic resistance compared to other lineages, diverse surface protein antigens, and a higher rate of recombination, a process by which bacteria can uptake and incorporate genetic material from its surroundings. Differences in vaccine use among sampled countries did not appear to be associated with the emergence of Clade II. We highlight the need to routine, representative sampling of bacterial isolates from diverse geographic areas and show the utility of genomic data in resolving epidemiological differences within a pathogen population.

evolutionary biology

pyseer: a comprehensive tool for microbial pangenome-wide association studies

SummaryGenome-wide association studies (GWAS) in microbes face different challenges to eukaryotes and have been addressed by a number of different methods. pyseer brings these techniques together in one package tailored to microbial GWAS, allows greater flexibility of the input data used, and adds new methods to interpret the association results.\n\nAvailability and Implementationpyseer is written in python and is freely available at https://github.com/mgalardini/pyseer, or can be installed through pip. Documentation and a tutorial are available at http://pyseer.readthedocs.io.\n\nContactjohn.lees@nyumc.org and marco@ebi.ac.uk\n\nSupplementary informationSupplementary data are available online.

bioinformatics

Global phylogenomics of multidrug-resistant Staphylococcus aureus sequence type 772: the Bengal Bay clone

The global spread of antimicrobial resistance has been well documented in Gram-negative bacteria and healthcare-associated epidemic pathogens, often emerging from regions with heavy antimicrobial use. However, the degree to which similar processes occur with Gram-positive bacteria in the community setting is less well understood. Here we demonstrate the recent origin and global spread from the Indian subcontinent of a multidrug resistant Staphylococcus aureus lineage, sequence type 772 (Bengal Bay clone). Short-term outbreaks occurred following intercontinental transmission, typically associated with travel and family contacts, but ongoing endemic transmission was uncommon. Instrumental in the emergence of a single dominant clade in the early 1990s was the acquisition of a multidrug resistance integrated plasmid that did not appear to incur a significant fitness cost. The Bengal Bay clone therefore combines the multidrug resistance of traditional healthcare-associated clones with the epidemiological and virulence potential of community-associated clones.

genomics

Pneumococcal vaccine impacts on the population genomics of non-typeable Haemophilus influenzae

Between 2008/09 and 2012/13 the molecular epidemiology of non-typeable Haemophilus influenzae (NTHi) carriage in children <5 years of age was determined; a period that included pneumococcal conjugate vaccine (PCV) 13 introduction. Significantly increased carriage in post-PCV13 years was observed and lineage-specific associations with S. pneumoniae were observed before and after PCV13 introduction. NTHi were characterised into eleven discrete, temporally stable lineages, congruent with current knowledge regarding the clonality of NTHi. This increase could not be linked to the expansion of a particular clone and demonstrates different dynamics to before PCV13 implementation during which time NTHi co-carried with vaccine serotype pneumococci.

microbiology

SuperDCA for genome-wide epistasis analysis

The potential for genome-wide modeling of epistasis has recently surfaced given the possibility of sequencing densely sampled populations and the emerging families of statistical interaction models. Direct coupling analysis (DCA) has earlier been shown to yield valuable predictions for single protein structures, and has recently been extended to genome-wide analysis of bacteria, identifying novel interactions in the co-evolution between resistance, virulence and core genome elements. However, earlier computational DCA methods have not been scalable to enable model fitting simultaneously to 104-105 polymorphisms, representing the amount of core genomic variation observed in analyses of many bacterial species. Here we introduce a novel inference method (SuperDCA) which employs a new scoring principle, efficient parallelization, optimization and filtering on phylogenetic information to achieve scalability for up to 105 polymorphisms. Using two large population samples of Streptococcus pneumoniae, we demonstrate the ability of SuperDCA to make additional significant biological findings about this major human pathogen. We also show that our method can uncover signals of selection that are not detectable by genome-wide association analysis, even though our analysis does not require phenotypic measurements. SuperDCA thus holds considerable potential in building understanding about numerous organisms at a systems biological level.\n\nAuthor SummaryRecent work has demonstrated the emerging potential in statistical genome-wide modeling to uncover co-selection and epistatic interactions between polymorphisms in bacterial chromosomes from densely sampled population data. Here we develop the Potts model based approach further into a fully mature computational method which can be applied to most existing bacterial population genomic data sets in a straightforward manner. Our advances are relying on more efficient parameter scoring, highly optimized and parallelized open source C++ code, which does not rely on the computation-intensive polymorphism subsampling approximations used earlier. By analyzing the two largest available population samples of Streptococcus pneumoniae (the pneumococcus), we highlight several biological discoveries related to the survival of the pneumococcus and co-evolution of penicillin-binding loci, which were not uncovered by the earlier analyses. Our method holds considerable potential for building understanding about numerous organisms at a systems biological level.

genomics

SeroBA: rapid high-throughput serotyping of Streptococcus pneumoniae from whole genome sequence data

Streptococcus pneumoniae is responsible for 240,000 - 460,000 deaths in children under 5 years of age each year. Accurate identification of pneumococcal serotypes is important for tracking the distribution and evolution of serotypes following the introduction of effective vaccines. Recent efforts have been made to infer serotypes directly from genomic data but current software approaches are limited and do not scale well. Here, we introduce a novel method, SeroBA, which uses a hybrid assembly and mapping approach. We compared SeroBA against real and simulated data and present results on the concordance and computational performance against a validation dataset, the robustness and scalability when analysing a large dataset, and the impact of varying the depth of coverage in the cps locus region on sequence-based serotyping. SeroBA can predict serotypes, by identifying the cps locus, directly from raw whole genome sequencing read data with 98% concordance using a k-mer based method, can process 10,000 samples in just over 1 day using a standard server and can call serotypes at a coverage as low as 10x. SeroBA is implemented in Python3 and is freely available under an open source GPLv3 license from: https://github.com/sanger-pathogens/seroba\n\nDATA SUMMARYO_LIThe reference genome Streptococcus pneumoniae ATCC 700669 is available from National Center for Biotechnology Information (NCBI) with the accession number: FM211187\nC_LIO_LISimulated paired end reads for experiment 2 have been deposited in FigShare: https://doi.org/10.6084/m9.figshare.5086054.v1\nC_LIO_LIAccession numbers for all other experiments are listed in Supplementary Table S1 and Supplementary Table S2.\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTThis article describes SeroBA, a k-mer based method for predicting the serotypes of Streptococcus pneumoniae from Whole Genome Sequencing (WGS) data. SeroBA can identify 92 serotypes and 2 subtypes with constant memory usage and low computational costs. We showed that SeroBA is able to reliably predict serotypes at a depth of coverage as low as 10x and is scalable to large datasets.

bioinformatics

PANINI: Pangenome Neighbor Identification for Bacterial Populations

The standard workhorse for genomic analysis of the evolution of bacterial populations is phylogenetic modelling of mutations in the core genome. However, in the current era of population genomics, a notable amount of information about evolutionary and transmission processes in diverse populations can be lost unless the accessory genome is also taken into consideration. Here we introduce PANINI, a computationally scalable method for identifying the neighbours for each isolate in a data set using unsupervised machine learning with stochastic neighbour embedding. PANINI is browser-based and integrates with the Microreact platform for rapid online visualisation and exploration of both core and accessory genome evolutionary signals together with relevant epidemiological, geographic, temporal and other metadata. Several case studies with single-and multi-clone pneumococcal populations are presented to demonstrate ability to identify biologically important signals from gene content data. PANINI is available at http://panini.wgsa.net/ and code at http://gitlab.com/cgps/panini

microbiology

Heterogeneity Among Estimates Of The Core Genome And Pan-Genome In Different Pneumococcal Populations

BackgroundUnderstanding the structure of a bacterial population is essential in order to understand bacterial evolution, or which genetic lineages cause disease, or the consequences of perturbations to the bacterial population. Estimating the core genome, the genes common to all or nearly all strains of a species, is an essential component of such analyses. The size and composition of the core genome varies by dataset, but our hypothesis was that variation between different collections of the same bacterial species should be minimal. To test this, the genome sequences of 3,121 pneumococci recovered from healthy individuals in Reykjavik (Iceland), Southampton (United Kingdom), Boston (USA) and Maela (Thailand) were analysed.\n\nResultsThe analyses revealed a supercore genome (genes shared by all 3,121 pneumococci) of only 303 genes, although 461 additional core genes were shared by pneumococci from Reykjavik, Southampton and Boston. Overall, the size and composition of the core genomes and pan-genomes among pneumococci recovered in Reykjavik, Southampton and Boston were very similar, but pneumococci from Maela were distinctly different. Inspection of the pan-genome of Maela pneumococci revealed several >25 Kb sequence regions that were homologous to genomic regions found in other bacterial species.\n\nConclusionsSome subsets of the global pneumococcal population are highly heterogeneous and thus our hypothesis was rejected. This is an essential point of consideration before generalising the findings from a single dataset to the wider pneumococcal population.

genomics

Methicillin resistant Staphylococcus aureus emerged long before the introduction of methicillin in to clinical practice

The spread of drug-resistant bacterial pathogens pose a major threat to global health. It is widely recognised that the widespread use of antibiotics has generated selective pressures that have driven the emergence of resistant strains. Methicillin-resistant Staphylococcus aureus (MRSA) was first observed in 1960, less than one year after the introduction of this second generation {beta}-lactam antibiotic into clinical practice. Epidemiological evidence has always suggested that resistance arose around this period, when the mecA gene encoding methicillin resistance carried on an SCCmec element, was horizontally transferred to an intrinsically sensitive strain of S. aureus. Whole genome sequencing a collection of the very first MRSA isolates allowed us to reconstruct the evolutionary history of the archetypal MRSA. Bayesian phylogenetic reconstruction was applied to infer the time point at which this early MRSA lineage arose and when SCCmec was acquired. MRSA emerged in the mid 1940s, following the acquisition of an ancestral type I SCCmec element, some fourteen years prior to the first therapeutic use of methicillin. Methicillin use was not the original driving factor in the evolution of MRSA as previously thought. Rather it was the widespread use of first generation {beta}-lactams such as penicillin in the years prior to the introduction of methicillin, which selected for S. aureus strains carrying the mecA determinant. Crucially this highlights how new drugs, introduced to circumvent known resistance mechanisms, can be rendered ineffective by unrecognised adaptations in the bacterial population due to the historic selective landscape created by the widespread use of other antibiotics.

microbiology

Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration

Streptococcus pneumoniae is a leading cause of invasive disease in infants, especially in low-income settings. Asymptomatic carriage in the nasopharynx is a prerequisite for disease, and the duration of carriage is an important consideration in modelling transmission dynamics and vaccine response. Existing studies of carriage duration variability are based at the serotype level only, and do not probe variation within lineages or fully quantify interactions with other environmental factors.\n\nHere we developed a model to calculate the duration of carriage episodes from longitudinal swab data. By combining these results with whole genome sequence data we estimate that pneumococcal genomic variation accounted for 63% of the phenotype variation, whereas host traits accounted for less than 5%. We further partitioned this heritability into both lineage and locus effects, and quantified the amount attributable to the largest sources of variation in carriage duration: serotype (17%), drug-resistance (9%) and other significant locus effects (7%). For the locus effects, a genome-wide association study identified 16 loci which may have an effect on carriage duration independent of serotype. Hits at a genome-wide level of significance were to prophage sequences, suggesting infection by such viruses substantially affects carriage duration.\n\nThese results show that both serotype and non-serotype specific effects alter carriage duration in infants and young children and are more important than other environmental factors such as host genetics. This has implications for models of pneumococcal competition and antibiotic resistance, and leads the way for the analysis of heritability of complex bacterial traits.\n\nSignificance statementOther than serotype, the genetic determinants of pneumococcal carriage duration are unknown. In this study we used longitudinal sampling to measure the duration of carriage in infants, and searched for any associated variation in the pan-genome. While we found that the pathogen genome explains most of the variability in duration, serotype did not fully account for this. Recent theoretical work has proposed the existence of alleles which alter carriage duration to explain the puzzle of continued coexistence of antibiotic-resistant and sensitive strains. Here we have shown that these alleles do exist in a natural population, and also identified candidates for the loci which fulfil this role. Together these findings have implications for future modelling of pneumococcal epidemiology and resistance.

genomics

Frequent recombination of pneumococcal capsule highlights future risks of emergence of novel serotypes.

Capsular diversity of Streptococcus pneumoniae constitutes a major obstacle in eliminating the pneumococcal disease. Such diversity is genetically encoded by almost 100 variants of the capsule polysaccharide locus (cps). However, the evolutionary dynamics of the capsule - the target of the currently used vaccines - remains not fully understood. Here, using genetic data from 4,469 bacterial isolates, we found cps to be an evolutionary hotspot with elevated substitution and recombination rates. These rates were a consequence of altered selection at this locus, supporting the hypothesis that the capsule has an increased potential to generate novel diversity compared to the rest of the genome. Analysis of twelve serogroups revealed their complex evolutionary history, which was principally driven by recombination with other serogroups and other streptococci. We observed significant variation in recombination rates between different serogroups. This variation could only be partially explained by the lineage-specific recombination rate, the remaining factors being likely driven by serogroup-specific ecology and epidemiology. Finally, we discovered two previously unobserved mosaic serotypes in the densely sampled collection from Mae La, Thailand, here termed 10X and 21X. Our results thus emphasise the strong adaptive potential of the bacterium by its ability to generate novel serotypes by recombination.

evolutionary biology