Search bioRxivSearch

Biology subjects

Croucher, N. J.

Publications and source records attributed to Croucher, N. J..

9 recordsLinked to original sources

Joint sequencing of human and pathogen genomes reveals the genetics of pneumococcal meningitis

Streptococcus pneumoniae is a common nasopharyngeal colonizer, but can also cause life-threatening invasive diseases such as empyema, bacteremia and meningitis. Genetic variation of host and pathogen is known to play a role in invasive pneumococcal disease, though to what extent is unknown. In a genome-wide association study of human and pathogen we show that human variation explains almost half of variation in susceptibility to pneumococcal meningitis and one-third of variation in severity, and identified variants in CCDC33 associated with susceptibility. Pneumococcal variation explained a large amount of invasive potential, but serotype explained only half of this variation. Newly developed methods identified pneumococcal genes involved in invasiveness including pspC and zmpD, and allowed a human-bacteria interaction analysis, finding associations between pneumococcal lineage and STK32C.

genomics

Fast and flexible bacterial genomic epidemiology with PopPUNK

The routine use of genomics for disease surveillance provides the opportunity for high-resolution bacterial epidemiology.\n\nHowever, current whole-genome clustering and multi-locus typing approaches do not fully exploit core and accessory genomic variation, and cannot both automatically identify, and subsequently expand, clusters of significantly-similar isolates in large datasets and across species.\n\nHere we describe PopPUNK (Population Partitioning Using Nucleotide K-mers; https://poppunk.readthedocs.io/en/latest/). software implementing scalable and expandable annotation- and alignment-free methods for population analysis and clustering.\n\nVariable-length k-mer comparisons are used to distinguish isolates divergence in shared sequence and gene content, which we demonstrate to be accurate over multiple orders of magnitude using both simulated data and real datasets from ten taxonomically-widespread species. Connections between closely-related isolates of the same strain are robustly identified, despite variation in the discontinuous pairwise distance distributions that reflects species diverse evolutionary patterns. PopPUNK can process 103-104 genomes as single batch, with minimal memory use and runtimes up to 200-fold faster than existing methods. Clusters of strains remain consistent as new batches of genomes are added, which is achieved without needing to re-analyse all genomes de novo.\n\nThis facilitates real-time surveillance with stable cluster naming and allows for outbreak detection using hundreds of genomes in minutes. Interactive visualisation and online publication is streamlined through automatic output of results to multiple platforms.\n\nPopPUNK has been designed as a flexible platform that addresses important issues with currently used whole-genome clustering and typing methods, and has potential uses across bacterial genetics and public health research.

genomics

Bayesian inference of ancestral dates on bacterial phylogenetic trees

The sequencing and comparative analysis of a collection of bacterial genomes from a single species or lineage of interest can lead to key insights into its evolution, ecology or epidemiology. The tool of choice for such a study is often to build a phylogenetic tree, and more specifically when possible a dated phylogeny, in which the dates of all common ancestors are estimated. Here we propose a new Bayesian methodology to construct dated phylogenies which is specifically designed for bacterial genomics. Unlike previous Bayesian methods aimed at building dated phylogenies, we consider that the phylogenetic relationships between the genomes have been previously evaluated using a standard phylogenetic method, which makes our methodology much faster and scalable. This two-steps approach also allows us to directly exploit existing phylogenetic methods that detect bacterial recombination, and therefore to account for the effect of recombination in the construction of a dated phylogeny. We analysed many simulated datasets in order to benchmark the performance of our approach in a wide range of situations. Furthermore, we present applications to three different real datasets from recent bacterial genomic studies. Our methodology is implemented in a R package called BactDating which is freely available for download at https://github.com/xavierdidelot/BactDating.

bioinformatics

Population genomics of pneumococcal carriage in Massachusetts children following PCV-13 introduction

BackgroundThe 13-valent pneumococcal conjugate vaccine (PCV-13) was introduced in the United States in 2010. Using a large pediatric carriage sample collected from shortly after the introduction of PCV-7 to several years after the introduction of PCV-13, we investigate alterations in the composition of the pneumococcal population following the introduction of PCV-13, evaluating the extent to which the post-vaccination non-vaccine type (NVT) population mirrors that from prior to vaccine introduction and the effect of PCV-13 on vaccine type lineages.\n\nMethods and FindingsDraft genome assemblies from 736 newly sequenced and 616 previously published pneumococcal carriages isolates from children in Massachusetts between 2001 and 2014 were analyzed. Isolates were classified into one of 22 sequence clusters (SCs) on the basis of their core genome sequence. We calculated the SC diversity for each sampling period as the probability that any two randomly drawn isolates from that period belong to different SCs. The sampling period immediately after the introduction of PCV-13 (2011) was found to have higher diversity than preceding (2007) or subsequent (2014) sampling periods (Simpsons D 2007: 0.915 95% CI [0.901, 0.929]; 2011: 0.935 [0.927, 0.942]; 2014: 0.912 [0.901, 0.923]). Amongst NVT isolates, we found the distribution of SCs in 2011 to be significantly different from that in 2007 or 2014 (Fishers Exact Test p=0.018, 0.0078), but did not find a difference comparing 2007 to 2014 (Fishers Exact Test p=0.24), indicating greater similarity between samples separated by a longer time period than between samples from closer time periods. We also found changes in the accessory gene content of the NVT population between 2007 and 2011 to have been reduced by 2014. Amongst the new serotypes targeted by PCV-13, four were present in our sample. The proportion of our sample composed of PCV-13-only vaccine serotypes 19A, 6C, and 7F decreased between 2007 and 2014, but no such reduction was seen for serotype 3. We did, however, observe differences in the genetic composition of the pre- and post-PCV-13 serotype 3 population. Our isolates were collected during discrete sampling periods from a small geographic area, which may limit the generalizability our findings.\n\nConclusionPneumococcal diversity increased immediately following the introduction of PCV-13, but subsequently returned to pre-vaccination levels. This is reflected in the distribution of NVT lineages, and, to a lesser extent, their accessory gene frequencies. As such, there may be a period during which the population is particularly disrupted by vaccination before returning to a more stable distribution. The persistence and shifting genetic composition of serotype 3 is a concern and warrants further investigation.

evolutionary biology

SuperDCA for genome-wide epistasis analysis

The potential for genome-wide modeling of epistasis has recently surfaced given the possibility of sequencing densely sampled populations and the emerging families of statistical interaction models. Direct coupling analysis (DCA) has earlier been shown to yield valuable predictions for single protein structures, and has recently been extended to genome-wide analysis of bacteria, identifying novel interactions in the co-evolution between resistance, virulence and core genome elements. However, earlier computational DCA methods have not been scalable to enable model fitting simultaneously to 104-105 polymorphisms, representing the amount of core genomic variation observed in analyses of many bacterial species. Here we introduce a novel inference method (SuperDCA) which employs a new scoring principle, efficient parallelization, optimization and filtering on phylogenetic information to achieve scalability for up to 105 polymorphisms. Using two large population samples of Streptococcus pneumoniae, we demonstrate the ability of SuperDCA to make additional significant biological findings about this major human pathogen. We also show that our method can uncover signals of selection that are not detectable by genome-wide association analysis, even though our analysis does not require phenotypic measurements. SuperDCA thus holds considerable potential in building understanding about numerous organisms at a systems biological level.\n\nAuthor SummaryRecent work has demonstrated the emerging potential in statistical genome-wide modeling to uncover co-selection and epistatic interactions between polymorphisms in bacterial chromosomes from densely sampled population data. Here we develop the Potts model based approach further into a fully mature computational method which can be applied to most existing bacterial population genomic data sets in a straightforward manner. Our advances are relying on more efficient parameter scoring, highly optimized and parallelized open source C++ code, which does not rely on the computation-intensive polymorphism subsampling approximations used earlier. By analyzing the two largest available population samples of Streptococcus pneumoniae (the pneumococcus), we highlight several biological discoveries related to the survival of the pneumococcus and co-evolution of penicillin-binding loci, which were not uncovered by the earlier analyses. Our method holds considerable potential for building understanding about numerous organisms at a systems biological level.

genomics

PANINI: Pangenome Neighbor Identification for Bacterial Populations

The standard workhorse for genomic analysis of the evolution of bacterial populations is phylogenetic modelling of mutations in the core genome. However, in the current era of population genomics, a notable amount of information about evolutionary and transmission processes in diverse populations can be lost unless the accessory genome is also taken into consideration. Here we introduce PANINI, a computationally scalable method for identifying the neighbours for each isolate in a data set using unsupervised machine learning with stochastic neighbour embedding. PANINI is browser-based and integrates with the Microreact platform for rapid online visualisation and exploration of both core and accessory genome evolutionary signals together with relevant epidemiological, geographic, temporal and other metadata. Several case studies with single-and multi-clone pneumococcal populations are presented to demonstrate ability to identify biologically important signals from gene content data. PANINI is available at http://panini.wgsa.net/ and code at http://gitlab.com/cgps/panini

microbiology

Phandango: an interactive viewer for bacterial population genomics.

SummaryFully exploiting the wealth of data in current bacterial population genomics datasets requires synthesising and integrating different types of analysis across millions of base pairs in hundreds or thousands of isolates. Current approaches often use static representations of phylogenetic, epidemiological, statistical and evolutionary analysis results that are difficult to relate to one another. Phandango is an interactive application running in a web browser allowing fast exploration of large-scale population genomics datasets combining the output from multiple genomic analysis methods in an intuitive and interactive manner.\n\nAvailabilityPhandango is a web application freely available for use at https://jameshadfield.github.io/phandango and includes a diverse collection of datasets as examples. Source code together with a detailed wiki page is available on GitHub at https://github.com/jameshadfield/phandango\n\nContactjh22@sanger.ac.uk, sh16@sanger.ac.uk

bioinformatics

Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration

Streptococcus pneumoniae is a leading cause of invasive disease in infants, especially in low-income settings. Asymptomatic carriage in the nasopharynx is a prerequisite for disease, and the duration of carriage is an important consideration in modelling transmission dynamics and vaccine response. Existing studies of carriage duration variability are based at the serotype level only, and do not probe variation within lineages or fully quantify interactions with other environmental factors.\n\nHere we developed a model to calculate the duration of carriage episodes from longitudinal swab data. By combining these results with whole genome sequence data we estimate that pneumococcal genomic variation accounted for 63% of the phenotype variation, whereas host traits accounted for less than 5%. We further partitioned this heritability into both lineage and locus effects, and quantified the amount attributable to the largest sources of variation in carriage duration: serotype (17%), drug-resistance (9%) and other significant locus effects (7%). For the locus effects, a genome-wide association study identified 16 loci which may have an effect on carriage duration independent of serotype. Hits at a genome-wide level of significance were to prophage sequences, suggesting infection by such viruses substantially affects carriage duration.\n\nThese results show that both serotype and non-serotype specific effects alter carriage duration in infants and young children and are more important than other environmental factors such as host genetics. This has implications for models of pneumococcal competition and antibiotic resistance, and leads the way for the analysis of heritability of complex bacterial traits.\n\nSignificance statementOther than serotype, the genetic determinants of pneumococcal carriage duration are unknown. In this study we used longitudinal sampling to measure the duration of carriage in infants, and searched for any associated variation in the pan-genome. While we found that the pathogen genome explains most of the variability in duration, serotype did not fully account for this. Recent theoretical work has proposed the existence of alleles which alter carriage duration to explain the puzzle of continued coexistence of antibiotic-resistant and sensitive strains. Here we have shown that these alleles do exist in a natural population, and also identified candidates for the loci which fulfil this role. Together these findings have implications for future modelling of pneumococcal epidemiology and resistance.

genomics

Easy and Accurate Reconstruction of Whole HIV Genomes from Short-Read Sequence Data

Next-generation sequencing has yet to be widely adopted for HIV. The difficulty of accurately reconstructing the consensus sequence of a quasispecies from reads (short fragments of DNA) in the presence of rapid between- and within-host evolution may have presented a barrier. In particular, mapping (aligning) reads to a reference sequence leads to biased loss of information; this bias can distort epidemiological and evolutionary conclusions. De novo assembly avoids this bias by effectively aligning the reads to themselves, producing a set of sequences called contigs. However contigs provide only a partial summary of the reads, misassembly may result in their having an incorrect structure, and no information is available at parts of the genome where contigs could not be assembled. To address these problems we developed the tool shiver to preprocess reads for quality and contamination, then map them to a reference tailored to the sample using corrected contigs supplemented with existing reference sequences. Run with two commands per sample, it can easily be used for large heterogeneous data sets. We use shiver to reconstruct the consensus sequence and minority variant information from paired-end short-read data produced with the Illumina platform, for 65 existing publicly available samples and 50 new samples. We show the systematic superiority of mapping to shivers constructed reference over mapping the same reads to the standard reference HXB2: an average of 29 bases per sample are called differently, of which 98.5% are supported by higher coverage. We also provide a practical guide to working with imperfect contigs.

bioinformatics