Search bioRxivSearch

Biology subjects

Jukka Corander

Publications and source records attributed to Jukka Corander.

7 recordsLinked to original sources

Interacting networks of resistance, virulence and core machinery genes identified by genome-wide epistasis analysis

Recent advances in the scale and diversity of population genomic datasets for bacteria now provide the potential for genome-wide patterns of co-evolution to be studied at the resolution of individual bases. The major human pathogen Streptococcus pneumoniae represents the first bacterial organism for which densely enough sampled population data became available for such an analysis. Here we describe a new statistical method, genomeDCA, which uses recent advances in computational structural biology to identify the polymorphic loci under the strongest co-evolutionary pressures. Genome data from over three thousand pneumococcal isolates identified 5,199 putative epistatic interactions between 1,936 sites. Over three-quarters of the links were between sites within the pbp2x, pbp1a and pbp2b genes, the sequences of which are critical in determining non-susceptibility to beta-lactam antibiotics. A network-based analysis found these genes were also coupled to that encoding dihydrofolate reductase, changes to which underlie trimethoprim resistance. Distinct from these resistance genes, a large network component of 384 protein coding sequences encompassed many genes critical in basic cellular functions, while another distinct component included genes associated with virulence. These results have the potential both to identify previously unsuspected protein-protein interactions, as well as genes making independent contributions to the same phenotype. This approach greatly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.\n\nAuthor SummaryEpistatic interactions between polymorphisms in DNA are recognized as important drivers of evolution in numerous organisms. Study of epistasis in bacteria has been hampered by the lack of both densely sampled population genomic data, suitable statistical models and powerful inference algorithms for extremely high-dimensional parameter spaces. We introduce the first model-based method for genome-wide epistasis analysis and use the largest available bacterial population genome data set on Streptococcus pneumoniae (the pneumococcus) to demonstrate its potential for biological discovery. Our approach reveals interacting networks of resistance, virulence and core machinery genes in the pneumococcus, which highlights putative candidates for novel drug targets. Our method significantly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.

Genetics

Analysis of recent and ancestral recombination reveals high-resolution population structure in Streptococcus pneumoniae

Prokaryotic evolution is affected by horizontal transfer of genetic material through recombination. Inference of an evolutionary tree of bacteria thus relies on accurate identification of the population genetic structure and recombination-derived mosaicism. Rapidly growing databases represent a challenge for computational methods to detect recombinations in bacterial genomes. We introduce a novel algorithm called fastGEAR which identifies lineages in diverse microbial alignments, and recombinations between them and from external origins. The algorithm detects both recent recombinations (affecting a few isolates) and ancestral recombinations between detected lineages (affecting entire lineages), thus providing insight into recombinations affecting deep branches of the phylogenetic tree. In sim-ulations, fastGEAR had comparable power to detect recent recombinations and outstanding power to detect the ancestral ones, compared to state-of-the-art methods, often with a fraction of computational cost. We demonstrate the utility of the method by analysing a collection of 616 whole-genomes of a recombinogenic pathogen Streptococcus pneumoniae, for which the method provided a high-resolution view of recombination across the genome. We examined in detail the penicillin-binding genes across the Streptococcus genus, demonstrating previously undetected genetic exchanges between different species at these three loci. Hence, fastGEAR can be readily applied to investigate mosaicism in bacterial genes across multiple species. Finally, fastGEAR correctly identified many known recombination hotspots and pointed to potential new ones. Matlab code and Linux/Windows executables are available at https://users.ics.aalto.fi/~pemartti/fastGEAR/

Bioinformatics

Monomorphic genotypes within a generalist lineage of Campylobacter jejuni show signs of global dispersion

The decreased costs of genome sequencing have increased capability to apply whole-genome sequence on epidemiological surveillance of zoonotic Campylobacter jejuni. However, knowledge about how genetically similar epidemiologically linked isolates can be is vital for correct application of this methodology. To address this issue in C. jejuni we investigated the spatial and temporal signals in the genomes of a major clonal complex and generalist lineage, ST-45 CC, by exploiting the population structure and genealogy and applying genome-wide association analysis of 340 isolates from across Europe collected over a wide time-range. The occurrence and strength of the geographical signal varied between sublineages and followed the clonal frame when present, while no evidence of a temporal signal was found. Certain sublineages of ST-45 CC formed discrete and genetically isolated clades to which geography and time had left only negligible traces in the genomes. We hypothesize that these ST-45 CC clades form globally expanded monomorphic clones possibly spread across Europe by migratory birds. In addition, we observed an incongruence between the genealogy of the strains and MLST typing, thereby challenging the existing clonal complex definition and use of a common MLST-based nomenclature for the ST-45 CC of C. jejuni.

Microbiology

Transcriptome remodeling contributes to epidemic disease caused by the human pathogen Streptococcus pyogenes

For over a century, a fundamental objective in infection biology research has been to understand the molecular processes contributing to the origin and perpetuation of epidemics. Divergent hypotheses have emerged concerning the extent to which environmental events or pathogen evolution dominates in these processes. Remarkably few studies bear on this important issue. Based on population pathogenomic analysis of 1200 Streptococcus pyogenes type emm89 infection isolates, we report that a series of horizontal gene transfer events produced a new pathogenic genotype with increased ability to cause infection, leading to an epidemic wave of disease on at least two continents. In the aggregate, these and other genetic changes substantially remodeled the transcriptomes of the evolved progeny, causing extensive differential expression of virulence genes and altered pathogen - host interaction, including enhanced immune evasion. Our findings delineate the precise molecular genetic changes that occurred and enhance our understanding of the evolutionary processes that contribute to the emergence and persistence of epidemically successful pathogen clones. The data have significant implications for understanding bacterial epidemics and translational research efforts to blunt their detrimental effects.\n\nImportanceThe confluence of studies of molecular events underlying pathogen strain emergence, evolutionary genetic processes mediating altered virulence, and epidemics is in its infancy. Although understanding these events is necessary to develop new or improved strategies to protect health, surprisingly few studies have addressed this issue, in particular at the comprehensive population genomic level. Herein we establish that substantial remodeling of the transcriptome of the human-specific pathogen Streptococcus pyogenes by horizontal gene flow and other evolutionary genetic changes is a central factor in precipitating and perpetuating epidemic disease. The data unambiguously show that the key outcome of these molecular events is evolution of a new, more virulent pathogenic genotype. Our findings provide new understanding of epidemic disease.

Evolutionary Biology

Sequence element enrichment analysis to determine the genetic basis of bacterial phenotypes

Bacterial genomes vary extensively in terms of both gene content and gene sequence - this plasticity hampers the use of traditional SNP-based methods for identifying all genetic associations with phenotypic variation. Here we introduce a computationally scalable and widely applicable statistical method (SEER) for the identification of sequence elements that are significantly enriched in a phenotype of interest. SEER is applicable to even tens of thousands of genomes by counting variable-length k-mers using a distributed string-mining algorithm. Robust options are provided for association analysis that also correct for the clonal population structure of bacteria. Using large collections of genomes of the major human pathogens Streptococcus pneumoniae and Streptococcus pyogenes, SEER identifies relevant previously characterised resistance determinants for several antibiotics and discovers potential novel factors related to the invasiveness of S. pyogenes. We thus demonstrate that our method can answer important biologically and medically relevant questions.

Genomics

The impact of host metapopulation structure on the population genetics of colonizing bacteria

Many key bacterial pathogens are frequently carried asymptomatically, and the emergence and spread of these opportunistic pathogens can be driven, or mitigated, via demographic changes within the host population. These inter-host transmission dynamics combine with basic evolutionary parameters such as rates of mutation and recombination, population size and selection, to shape the genetic diversity within bacterial populations. Whilst many studies have focused on how molecular processes underpin bacterial population structure, the impact of host migration and the connectivity of the local populations has received far less attention. A stochastic neutral model incorporating heightened local transmission has been previously shown to fit closely with genetic data for several bacterial species. However, this model did not incorporate transmission limiting population stratification, nor the possibility of migration of strains between subpopulations, which we address here by presenting an extended model. The model captures the observed population patterns for the common nosocomial pathogens Staphylococcus epidermidis and Enterococcus faecalis, while Staphylococcus aureus and Enterococcus faecium display deviations attributable to adaptation. It is demonstrated analytically and numerically that expected strain relatedness may either increase or decrease as a function of increasing migration rate between subpopulations, being a complex function of the rate at which microepidemics occur in the metapopulation. Moreover, it is shown that in a structured population markedly different rates of evolution may lead to indistinguishable patterns of relatedness among bacterial strains; caution is thus required when drawing evolution inference in these cases.

Genetics

On the identifiability of transmission dynamic models for infectious diseases

Understanding the transmission dynamics of infectious diseases is important for both biological research and public health applications. It has been widely demonstrated that statistical modeling provides a firm basis for inferring relevant epidemiological quantities from incidence and molecular data. However, the complexity of transmission dynamic models causes two challenges: Firstly, the likelihood function of the models is generally not computable and computationally intensive simulation-based inference methods need to be employed. Secondly, the model may not be fully identifiable from the available data. While the first difficulty can be tackled by computational and algorithmic advances, the second obstacle is more fundamental. Identifiability issues may lead to inferences which are more driven by the prior assumptions than the data themselves. We here consider a popular and relatively simple, yet analytically intractable model for the spread of tuberculosis based on classical IS6110 fingerprinting data. We report on the identifiability of the model, presenting also some methodological advances regarding the inference. Using likelihood approximations, it is shown that the reproductive value cannot be identified from the data available and that the posterior distributions obtained in previous work have likely been substantially dominated by the assumed prior distribution. Further, we show that the inferences are influenced by the assumed infectious population size which has generally been kept fixed in previous work. We demonstrate that the infectious population size can be inferred if the remaining epidemiological parameters are already known with sufficient precision.

Bioinformatics