Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Solving the influence maximization problem reveals regulatory organization of the yeast cell cycle.

The Influence Maximization Problem (IMP) aims to discover the set of nodes with the greatest influence on network dynamics. The problem has previously been applied in epidemiology and social network analysis. Here, we demonstrate the application to cell cycle regulatory network analysis of Saccharomyces cerevisiae.\n\nFundamentally, gene regulation is linked to the flow of information. Therefore, our implementation of the IMP was framed as an information theoretic problem on a diffusion network. Utilizing all regulatory edges from YeastMine, gene expression dynamics were encoded as edge weights using a variant of time lagged transfer entropy, a method for quantifying information transfer between variables. Influence, for a particular number of sources, was measured using a diffusion model based on Markov chains with absorbing states. By maximizing over different numbers of sources, an influence ranking on genes was produced.\n\nThe influence ranking was compared to other metrics of network centrality. Although top genes from each centrality ranking contained well-known cell cycle regulators, there was little agreement and no clear winner. However, it was found that influential genes tend to directly regulate or sit upstream of genes ranked by other centrality measures. This is quantified by computing node reachability between gene sets; on average, 59% of central genes can be reached when starting from the influential set, compared to 7% of influential genes when starting at another centrality measure.\n\nThe influential nodes act as critical sources of information flow, potentially having a large impact on the state of the network. Biological events that affect influential nodes and thereby affect information flow could have a strong effect on network dynamics, potentially leading to disease.\n\nCode and example data can be found at: https://github.com/Gibbsdavidl/miergolf\n\nAuthor SummaryThe Influence Maximization Problem (IMP) is general and is applied in fields such as epidemiology, social network analysis, and as shown here, biological network analysis. The aim is to discover the set of regulatory genes with the greatest influence in the network dynamics. As gene regulation, fundamentally, is about the flow of information, the IMP was framed as an information theoretic problem. Dynamics were encoded as edge weights using time lagged transfer entropy, a quantity that defines information transfer across variables. The information flow was accomplished using a diffusion model based on Markov chains with absorbing states. Ant optimization was applied to solve the subset selection problem, recovering the most influential nodes.The influential nodes act as critical sources of information flow, potentially affecting the network state. Biological events that impact the influential nodes and thereby affecting normal information flow, could have a strong effect on the network, potentially leading to disease.

Bioinformatics

Identifying genetic variants that affect viability in large cohorts

A number of open questions in human evolutionary genetics would become tractable if we were able to directly measure evolutionary fitness. As a step towards this goal, we developed a method to examine whether individual genetic variants, or sets of genetic variants, currently influence viability. The approach consists in testing whether the frequency of an allele varies across ages, accounting for variation in ancestry. We applied it to the Genetic Epidemiology Research on Aging (GERA) cohort and to the parents of participants in the UK Biobank. Across the genome, we find only a few common variants with large effects on age-specific mortality: tagging the APOE {varepsilon}4 allele and near CHRNA3. These results suggest that when large, even late onset effects are kept at low frequency by purifying selection. Testing viability effects of sets of genetic variants that jointly influence one of 42 traits, we detect a number of strong signals. In participants of the UK Biobank study of British ancestry, we find that variants that delay puberty timing are enriched in longer-lived parents (P~6x10-6 for fathers and P~2x10-3 for mothers), consistent with epidemiological studies. Similarly, in mothers, variants associated with later age at first birth are associated with a longer lifespan (P~1x10-3). Signals are also observed for variants influencing cholesterol levels, risk of coronary artery disease, body mass index, as well as risk of asthma. These signals exhibit consistent effects in the GERA cohort and among participants of the UK Biobank of non-British ancestry. Moreover, we see marked differences between males and females, most notably at the CHRNA3 locus, and variants associated with risk of coronary artery disease and cholesterol levels. Beyond our findings, the analysis serves as a proof of principle for how upcoming biomedical datasets can be used to learn about selection effects in contemporary humans.

evolutionary biology

Estimation of sub-epidemic dynamics by means of Sequential Monte Carlo Approximate Bayesian Computation: an application to the Swiss HIV Cohort Study

Our ability to accurately infer transmission patterns of infectious diseases is critical to monitor both their spread and the efficacy of public health policies. The use of phylogenetic methods for the reconstruction of viral ancestral relationships has garnered increasing interest, particularly in the characterization of HIV epidemics and sub-epidemics. In the case of this virus, the Swiss HIV Cohort Study (SHCS) contains a wide breadth of genomic data that have been widely used as a means of applying such methods. However, current approaches for quantifying the epidemiological dynamics of diseases are computationally intensive, and fail to scale well with this magnitude of data. To address this issue, we re-implement an Approximate Bayesian Computation (ABC) approach based on sequential Monte Carlo (SMC). By means of simulations, we demonstrate that our implementation is capable of inferring key epidemiological parameters of the Swiss HIV epidemic accurately, and that sampling intensity has no significant effect on the accuracy of our estimates. Applied to a subset of HIV sequences from the SHCS, we show that we can distinguish sub-epidemics that are circulating in culturally distinct Swiss regions. Given these findings, we propose that ABC-SMC samplers will allow us to evaluate the impact of new public health policies, such as the implementation of a needle exchange program in the case of HIV, based on genetic data sampled before and after the implementation of a new policy.

evolutionary biology

Biased phylodynamic inferences from analysing clusters of viral sequences

Phylogenetic methods are being increasingly used to help understand the transmission dynamics of measurably evolving viruses, including HIV. Clusters of highly similar sequences are often observed, which appear to follow a power law behaviour, with a small number of very large clusters. These clusters may help to identify subpopulations in an epidemic, and inform where intervention strategies should be implemented. However, clustering of samples does not necessarily imply the presence of a subpopulation with high transmission rates, as groups of closely related viruses can also occur due to non-epidemiological effects such as over-sampling. It is important to ensure that observed phylogenetic clustering reflects true heterogeneity in the transmitting population, and is not being driven by non-epidemiological effects.\n\nWe quantify the effect of using a falsely identified transmission cluster of sequences to estimate phylodynamic parameters including the effective population size and exponential growth rate. Our simulation studies show that taking the maximum size cluster to re-estimate parameters from trees simulated under a randomly mixing, constant population size coalescent process systematically underestimates the overall effective population size. In addition, the transmission cluster wrongly resembles an exponential or logistic growth model 95% of the time. We also illustrate the consequences of false clusters in exponentially growing coalescent and birth-death trees, where again, the growth rate is skewed upwards. This has clear implications for identifying clusters in large viral databases, where a false cluster could result in wasted intervention resources.

genetics

Pathogen mitigation in an Ecuadorian potato seed system: Insights from network analysis

Seed systems have an important role in the distribution of high quality seed and improved varieties. The structure of seed networks also helps to determine the epidemiological risk for seedborne disease. We present a new method for evaluating the epidemiological role of nodes in seed networks, and apply it to a regional potato farmer consortium (CONPAPA) in Ecuador. We surveyed farmers to estimate the structure of networks of farmer seed tuber and ware potato transactions, and farmer information sources about pest and disease management. Then we simulated pathogen spread through seed transaction networks to identify priority nodes for disease detection. The likelihood of pathogen establishment was weighted based on the quality and/or quantity of information sources about disease management. CONPAPA staff and facilities, a market, and certain farms are priorities for disease management interventions, such as training, monitoring and variety dissemination. Advice from agrochemical store staff was common but assessed as significantly less reliable. Farmer access to information (reported number and quality of sources) was similar for both genders. Women had a smaller amount of the market share for seed-tubers and ware potato, however. Understanding seed system networks provides input for scenario analyses to evaluate potential system improvements.

ecology

Evolution and manipulation of vector host choice

The transmission of many animal and plant diseases relies on the behavior of arthropod vectors. In particular, the choice to feed on either infected or uninfected hosts can dramatically affect the epidemiology of vector-borne diseases. I develop an epidemiological model to explore the impact of host choice behavior on the dynamics of these diseases and to examine selection acting on vector behavior, but also on pathogen manipulation of this behavior. This model identifies multiple evolutionary conflicts over the control of this behavior and generates testable predictions under different scenarios. In general, the vector should evolve the ability to avoid infected hosts. However, if the vector behavior is under the control of the pathogen, uninfected vectors should prefer infected hosts while infected vectors should seek uninfected hosts. But some mechanistic constraints on pathogen manipulation ability may alter these predictions. These theoretical results are discussed in the light of observed behavioral patterns obtained on a diverse range of vector-borne diseases. These patterns confirm that several pathogens have evolved conditional behavioral manipulation strategies of their vector species. Other pathogens, however, seem unable to evolve such complex conditional strategies. Contrasting the behavior of infected and uninfected vectors may thus help reveal mechanistic constraints acting on the evolution of the manipulation of vector behavior.

evolutionary biology

Co-evolution of virulence and immunosuppression through multiple infections

Many components of host-parasite interactions have been shown to affect the way virulence (i.e., parasite-induced harm to the host) evolves. However, coevolution of multiple parasite traits is often neglected. We explore how an immunosuppressive mechanism of parasites affects and co-evolves with virulence through multiple infections. Applying the adaptive dynamics framework to epidemiological models with coinfection, we show that immunosuppression is a double-edged-sword for the evolution of virulence. On one hand, it amplifies the adaptive benefit of virulence by increasing the abundance of coinfections through epidemiological feedbacks. On the other hand, immunosuppression hinders host recovery, prolonging the duration of infection and elevating the cost of killing the host. The balance between the cost and benefit of immunosuppression varies across different background mortality rates of hosts. In addition, we find that immunosuppression evolution is influenced considerably by the precise trade-off shape determining the effect of immunosuppression on host recovery and susceptibility to further infection. These results demonstrate that the evolution of virulence is shaped by immunosuppression while highlighting that the evolution of immune evasion mechanisms deserves further research attention.

evolutionary biology

Model-based analysis of experimental hut data elucidates multifaceted effects of a volatile chemical on Aedes aegypti mosquitoes

BackgroundInsecticides used against Aedes aegypti and other disease vectors can elicit a multitude of dose-dependent effects on behavioral and bionomic traits. Estimating the potential epidemiological impact of a product requires thorough understanding of these effects and their interplay at different dosages. Volatile spatial repellent (SR) products come with an additional layer of complexity due to the potential for movement of affected mosquitoes or volatile particles of the product beyond the treated house. Here, we propose a statistical inference framework for estimating these nuanced effects of volatile SRs.\n\nMethodsWe fitted a continuous-time Markov chain model in a Bayesian framework to mark-release-recapture (MRR) data from an experimental hut study conducted in Iquitos, Peru. We estimated the effects of two dosages of transfluthrin on Ae. aegypti behaviors associated with human-vector contact: repellency, exiting, and knockdown in the treated space and in \"downstream\" adjacent huts. We validated the framework using simulated data.\n\nResultsThe odds of a female Ae. aegypti being repelled from a treated hut (HT) increased at both dosages (low dosage: odds = 1.64, 95% highest density interval (HDI) = 1.30-2.09; high dosage: odds = 1.35, HDI = 1.04-1.67). The relative risk of exiting from the treated hut was reduced (low: RR = 0.70, HDI = 0.62-1.09; high: RR = 0.70, HDI = 0.40-1.06), with this effect carrying over to untreated spaces as far as two huts away from the treated hut (H2) (low: RR = 0.79, HDI = 0.59-1.01; high: RR = 0.66, HDI = 0.50-0.87). Knockdown rates were increased in both treated and downstream huts, particularly under high dosage (HT: RR = 8.37, HDI = 2.11-17.35; H1: RR = 1.39, HDI = 0.52-2.69; H2: RR = 2.22, HDI = 0.96-3.86).\n\nConclusionsOur statistical inference framework is effective at elucidating multiple effects of volatile chemicals used in SR products, as well as their downstream effects. This framework provides a powerful tool for early selection of candidate SR product formulations worth advancing to costlier epidemiological trials, which are ultimately necessary for proof of concept of public health value and subsequent formal endorsement by health authorities.

ecology

Vibrio cholerae genomic diversity within and between patients

Cholera is a severe, waterborne diarrheal disease caused by toxin-producing strains of the bacterium Vibrio cholerae. Comparative genomics has revealed \"waves\" of cholera transmission and evolution, in which clones are successively replaced over decades and centuries. However, the extent of V. cholerae genetic diversity within an epidemic or even within an individual patient is poorly understood. Here, we characterized V. cholerae genomic diversity at a micro-epidemiological level within and between individual patients from Bangladesh and Haiti. To capture within-patient diversity, we isolated multiple (8 to 20) V. cholerae colonies from each of eight patients, sequenced their genomes and identified point mutations and gene gain/loss events. We found limited but detectable diversity at the level of point mutations within hosts (zero to three single nucleotide variants within each patient), and comparatively higher gene content variation within hosts (at least one gain/loss event per patient, and up to 103 events in one patient). Much of the gene content variation appeared to be due to gain and loss of phage and plasmids within the V. cholerae population, with occasional exchanges between V. cholerae and other members of the gut microbiota. We also show that certain intra-host variants have phenotypic consequences. For example, the acquisition of a Bacteroides plasmid and nonsynonymous mutations in a sensor histidine kinase gene both reduced biofilm formation, an important trait for environmental survival. Together, our results show that V. cholerae is measurably evolving within patients, with possible implications for disease outcomes and transmission dynamics.\n\nAuthor SummaryVibrio cholerae is the etiological agent of cholera, a severe diarrheal disease endemic to Bangladesh and responsible for global outbreaks, including one ongoing in Haiti. Certain bacterial pathogens can evolve and diversify within the human host, often altering virulence and antibiotic resistance. However, most examples of within-host evolution have come from chronic infections, in which the pathogen has sufficient time to mutate and diversify, and little attention has been paid to more acute infections such as the one caused by V. cholerae. The goal of this study was to measure the extent of within-host evolution of V. cholerae within individual infected patients. By sequencing multiple bacterial isolates_from each of eight patients from Bangladesh and Haiti, we found that cholera patients can harbor a diverse population of V. cholerae. As expected for an acute infection, this diversity is limited, ranging from zero to three point mutations (single nucleotide variants) per patient. However, gene gain/loss events are more prevalent than point mutations, occurring in every single patient, and sometimes involving the transfer of dozens of genes on plasmids. Even if rare, point mutations and gene gain/loss events may be maintained by natural selection, and can alter clinically-and environmentally-relevant phenotypes such as biofilm formation. Therefore, within-patient evolution has the potential to impact clinical and epidemiological outcomes. Together, our results demonstrate that within-patient evolution may be a general feature of both acute and chronic infections, and that gene gain/loss may be an important but under-appreciated feature of within-host evolution.

microbiology

Modulation of host learning in Aedes aegypti mosquitoes

How mosquitoes determine which individuals to bite has important epidemiological consequences. This choice is not random; most mosquitoes specialize in one or a few vertebrate host species, and some individuals in a host population are preferred over others. Here we show that aversive olfactory learning contributes to mosquito preference both between and within host species. Combined electrophysiological and behavioural recordings from tethered flying mosquitoes demonstrated that these odours evoke changes in both behaviour and antennal lobe (AL) neuronal responses. Using electrophysiological and behavioural approaches, and CRISPR gene editing, we demonstrate that dopamine plays a critical role in aversive olfactory learning and modulating odour-evoked responses in AL neurons. Collectively, these results provide the first experimental evidence that olfactory learning in mosquitoes can play an epidemiological role.

neuroscience

Evolution-informed forecasting of seasonal influenza A (H3N2)

Inter-pandemic or seasonal influenza exacts an enormous annual burden both in terms of human health and economic impact. Incidence prediction ahead of season remains a challenge largely because of the virus antigenic evolution. We propose here a forecasting approach that incorporates evolutionary change into a mechanistic epidemiological model. The proposed models are simple enough that their parameters can be estimated from retrospective surveillance data. These models link amino-acid sequences of hemagglutinin epitopes with a transmission model for seasonal H3N2 influenza, also informed by H1N1 levels. With a monthly time series of H3N2 incidence in the United States over 10 years, we demonstrate the feasibility of prediction ahead of season and an accurate real-time forecast for the 2016/2017 influenza season.\n\nSUMMARYSkillful forecasting of seasonal (H3N2) influenza incidence ahead of the season is shown to be possible by means of a transmission model that explicitly tracks evolutionary change in the virus, integrating information from both epidemiological surveillance and readily available genetic sequences.

systems biology

Polygenic risk scores applied to a single cohort reveal pleiotropy among hundreds of human phenotypes

BackgroundThere is now convincing evidence that pleiotropy across the genome contributes to the correlation between human traits and comorbidity of diseases. The recent availability of genome-wide association study (GWAS) results have made the polygenic risk score (PRS) approach a powerful way to perform genetic prediction and identify genetic overlap among phenotypes.\n\nMethods and findingsHere we use the PRS method to assess evidence for shared genetic aetiology across hundreds of traits within a single epidemiological study - the Northern Finland Birth Cohort 1966 (NFBC1966). We replicate numerous recent findings, such as a genetic association between Alzheimers disease and lipid levels, while the depth of phenotyping in the NFBC1966 highlights a range of novel significant genetic associations between traits.\n\nConclusionsThis study illustrates the power in taking a hypothesis-free approach to the study of shared genetic aetiology between human traits and diseases. It also demonstrates the potential of the PRS method to provide important biological insights using only a single well-phenotyped epidemiological study of moderate sample size (~5k), with important advantages over evaluating genetic correlations from GWAS summary statistics only.

genetics

From Cohorts to Molecules: Adverse Impacts of Endocrine Disrupting Mixtures

Convergent evidence associates endocrine disrupting chemicals (EDCs) with major, increasingly-prevalent human disorders. Regulation requires elucidation of EDC-triggered molecular events causally linked to adverse health outcomes, but two factors limit their identification. First, experiments frequently use individual chemicals, whereas real life entails simultaneous exposure to multiple EDCs. Second, population-based and experimental studies are seldom integrated. This drawback was exacerbated until recently by lack of physiopathologically meaningful human experimental systems that link epidemiological data with results from model organisms.\n\nWe developed a novel approach, integrating epidemiological with experimental evidence. Starting from 1,874 mother-child pairs we identified mixtures of chemicals, measured during early pregnancy, associated with language delay or low-birth weight in offspring. These mixtures were then tested on multiple complementary in vitro and in vivo models. We demonstrate that each EDC mixture, at levels found in pregnant women, disrupts hormone-regulated and disease-relevant gene regulatory networks at both the cellular and organismal scale.

molecular biology

ITHANET: Information and database community portal for haemoglobinopathies

Haemoglobinopathies are the commonest monogenic diseases, with millions of carriers and patients worldwide. Online resources for haemoglobinopathies are largely divided into specialised sites catering for patients, researchers and clinicians separately. However, the severity, ubiquity and surprising genetic complexity of the haemoglobinopathies call for an integrated website to serve as a free and comprehensive repository and tool for patients, scientists and health professionals alike. This paper presents the ITHANET community portal, an expanding resource for clinicians and researchers dealing with haemoglobinopathies. It integrates information on news, events, publications, clinical trials and haemoglobinopathy-related organisations and experts and, most importantly, databases of variations, epidemiology and diagnostic and clinical data. Specifically, ITHANET provides annotation for 2690 haemoglobinopathy-related variations, epidemiological data for more than 180 countries and information for more than 600 HPLC diagnostic reports. The ITHANET portal accepts and incorporates contributions to its content by local experts from any country in the world and is freely available for the public at http://www.ithanet.eu.

bioinformatics

De Novo Mutations Resolve Disease Transmission Pathways in Clonal Malaria

Detecting de novo mutations in viral and bacterial pathogens enables researchers to reconstruct detailed networks of disease transmission and is a key technique in genomic epidemiology. However these techniques have not yet been applied to the malaria parasite, Plasmodium falciparum, in which a larger genome, slower generation times, and a complex life cycle make them difficult to implement. Here we demonstrate the viability of de novo mutation studies in P. falciparum for the first time. Using a set of clinical samples and novel methods of sequencing, library preparation, and genotyping, we have genotyped low-complexity regions of the genome with a high degree of accuracy. Despite its slower evolutionary rate compared to bacterial or viral species, de novo mutation can be detected in P. falciparum across timescales of just 1-2 years and evolutionary rates in low-complexity regions of the genome can be up to twice that detected in the rest of the genome. The increased mutation rate allows the identification of separate clade expansions that cannot be found using previous genomic epidemiology approaches and could be a crucial tool for mapping residual transmission patterns in disease elimination campaigns and reintroduction scenarios.

genomics

Resolving outbreak dynamics using Approximate Bayesian Computation for stochastic birth-death models

Earlier research has suggested that Approximate Bayesian Computation (ABC) makes it possible to fit simulator-based intractable birth-death models to investigate communicable disease outbreak dynamics with accuracy comparable to that of exact Bayesian methods. However, recent findings have indicated that key parameters such as the reproductive number R may remain poorly identifiable. Here we show that the identifiability issue can be resolved by taking into account disease-specific characteristics of the transmission process in closer detail. Using tuberculosis (TB) in the San Francisco Bay area as a case-study, we consider the situation where the genotype data are generated as a mixture of three stochastic processes, each with their distinct dynamics and clear epidemiological interpretation.\n\nThe ABC inference yields stable and accurate posterior inferences about outbreak dynamics from aggregated annual case data with genotype information. We also show that under the proposed model, the infectious population size can be reliably inferred from the data. The estimate is approximately two orders of magnitude smaller compared to assumptions made in the earlier ABC studies, and is much better aligned with epidemiological knowledge about active TB prevalence. Similarly, the reproductive number R related to the primary underlying transmission process is estimated to be nearly three-fold compared with the previous estimates, which has a substantial impact on the interpretation of the fitted outbreak model.

bioinformatics

Outbreak of invasive wound mucormycosis in a burn unit due to multiple strains of Mucor circinelloides f. circinelloides resolved by whole genome sequencing

Mucorales are ubiquitous environmental molds responsible for mucormycosis in diabetic, immunocompromised, and severely burned patients. Small outbreaks of invasive wound mucormycosis (IWM) have already been reported in burn units without extensive microbiological investigations. We faced an outbreak of IWM in our center and investigated the clinical isolates with whole genome sequencing (WGS) analysis.\n\nWe analyzed M. circinelloides isolates from patients in our burn unit (BU1) together with non-outbreak isolates from burn unit 2 (BU2, Paris area) and from France over a two-year period (2013-2015). For each isolate, WGS and a de novo genome assembly was performed from read data extracted from the aligned contig sequences of the reference genome (1006PhL).\n\nA total of 21 isolates were sequenced including 14 isolates from six BU1 patients. Phylogenetic classification showed that the clinical isolates clustered in four highly divergent clades. Clade1 contained at least one of the strains from the six epidemiologically-linked BU1 patients. The clinical isolates seemed specific to each patient. Two patients were infected with more than two strains from different clades suggesting that an environmental reservoir of clonally unrelated isolates was the source of contamination. Only two patients shared one strain in BU1, suggesting direct transmission or contamination with the same environmental source.\n\nWGS coupled with precise epidemiological data and analysis of several isolates per patients revealed in our study a complex situation with both potential cross-transmission and multiple contaminations with a heterogeneous pool of strains from a cryptic environmental reservoir.\n\nImportanceInvasive wound mucormycosis (IWM) is a severe infection due to the environmental molds belonging to the order Mucorales. Severely burned patients are particularly at risk for IWM. Here, we used Whole Genome Sequencing (WGS) analysis to resolve an outbreak of IWM due to Mucor circinelloides that occurred in our hospital (BU1). We sequenced 21 clinical isolates, including 14 from BU1 and 7 unrelated isolates, and compared them to the reference genome (1006PhL). This analysis revealed that the outbreak was mainly due to multiple strains that seemed patient-specific, suggesting that the patients were more likely infected from a pool of diverse strains from the environment rather than from direct transmission between the patients. This study revealed the complexity of a Mucorales outbreak in the settings of IWM in burn patients, which has been highlighted based on whole genome sequencing and careful sampling.

microbiology

Gene composition as a potential barrier to large recombinations in the bacterial pathogen Klebsiella pneumoniae

Klebsiella pneumoniae (Kp) is one of the most important nosocomial pathogens world-wide, being responsible for frequent hospital outbreaks and causing sepsis and multi-organ infections with a high mortality rate and frequent hospital outbreaks. The most prevalent and widely disseminated lineage of K. pneumoniae is clonal group 258 (CG258), which includes the highly resistant \"high-risk\" genotypes ST258 and ST11. Recent studies revealed that very large recombination events have occurred during the recent emergence of Kp lineages. A striking example is provided by ST258, which has undergone a recombination event that replaced over 1 Mb of the genome with DNA from an unrelated Kp donor. Although several examples of this phenomenon have been documented in Kp and other bacterial species, the significance of these very large recombination events for the emergence of either hyper-virulent or resistant clones remains unclear. Here we present an analysis of 834 Kp genomes that provides data on the frequency of these very large recombination events (defined as those involving >100Kb), their distribution within the genome, and the dynamics of gene flow within the Kp population. We note that very large recombination events occur frequently, and in multiple lineages, and that the majority of recombinational exchanges are clustered within two overlapping genomic regions, which result to be involved by recombination events with different frequencies. Our results also indicate that certain non-CG258 lineages are more likely to act as donors to CG258 recipients than others. Furthermore, comparison of gene content in CG258 and non-CG258 strains agrees with this pattern, suggesting that the success of a large recombination depends on gene composition in the exchanged genomic portion.\n\nAuthor SummaryKlebsiella pneumoniae (Kp) is an opportunistic bacterial pathogen, a major cause of deadly infections and outbreaks in hospitals worldwide. This bacterium is able to exchange large genomic portions (up to a fourth of the entire genome) within a single recombination event. Indeed, the most epidemiologically important Kp clone, is actually a hybrid which emerged after a > 1Mb recombination event. In this work, we investigated how recombinations affected the evolution of the most studied Kp Clonal Group, CG258. We found that large recombinations occurred frequently during Kp evolution, and occurred preferentially in a well-delimited genomic region. Furthermore, we found that four epidemiologically important clones emerged after large recombinations. We identified the donors of several large recombinations: despite many Kp lineages acted as donors during CG258 evolution, two of them have been involved more frequently. We hypothesize that the observed pattern of donors-recipients in recombinations, and the presence of a large recombinogenic region in Kp genome, could be related to gene composition. Indeed, genomic analyses showed a pattern compatible with this hypothesis, suggesting that gene content can represent a main factor in the success of a large recombination.

evolutionary biology