Search bioRxivSearch

Biology subjects

Colijn, C.

Publications and source records attributed to Colijn, C..

6 recordsLinked to original sources

Beyond the SNP threshold: identifying outbreak clusters using inferred transmissions

Whole genome sequencing (WGS) is increasingly used to aid in understanding pathogen transmission [1]. Very often the number of single nucleotide polymorphisms (SNPs) separating isolates collected during an epidemiological study are used to identify sets of cases that are potentially linked by direct transmission. However, there is little agreement in the literature as to what an appropriate SNP cut-off threshold should be, or indeed whether a simple SNP threshold is appropriate for identifying sets of isolates to be treated as \"transmission clusters\". The SNP thresholds that have been adopted for inferring transmission vary widely even for one pathogen. As an alternative to reliance on a strict SNP threshold, we suggest that the key inferential target when studying the spread of an infectious disease is the number of transmission events separating cases. Here we describe a new framework for deciding whether two pathogen genomes should be considered as part of the same transmission cluster, based jointly on the number of SNP differences and the length of time over which those differences have accumulated. Our approach allows us to probabilistically characterize the number of inferred transmission events that separate cases. We show how this framework can be modified to consider variable mutation rates across the genome (e.g. SNPs associated with drug resistance) and we indicate how the methodology can be extended to incorporate epidemiological data such as spatial proximity. We use recent data collected from tuberculosis studies from British Columbia, Canada and the Republic of Moldova to apply and compare our clustering method to the SNP threshold approach. In the British Columbia data, different cases break off from the main clusters as cut-off thresholds are lowered; the transmission-based method obtains slightly different clusters than the SNP cut-offs. For the Moldova data, straightforward application of the methods shows no appreciable difference, but when we take into account the fact that resistance conferring sites likely do not follow the same mutation clock as most sites due to selection, the transmission-based approach differs from the SNP cut-off method. Outbreak simulations confirm that our transmission based method is at least as good at identifying direct transmissions as a SNP cut-off. We conclude that the new method is a promising step towards establishing a more robust identification of outbreaks.

genomics

Comparing phylogenetic trees according to tip label categories

Trees that illustrate patterns of ancestry and evolution are a central tool in many areas of biology. Comparing evolutionary trees to each other has widespread applications in comparing the evolutionary stories told by different sources of data, assessing the quality of inference methods, and highlighting areas where patterns of ancestry are uncertain. While these tasks are complicated by the fact that trees are high-dimensional structures encoding a large amount of information, there are a number of metrics suitable for comparing evolutionary trees whose tips have the same set of unique labels. There are also metrics for comparing trees where there is no relationship between their labels: in unlabelled tree metrics the tree shapes are compared without reference to the tip labels.\n\nIn many interesting applications, however, the taxa present in two or more trees are related but not identical, and it is informative to compare the trees whilst retaining information about their tips relationships. We present methods for comparing trees whose labels belong to a pre-defined set of categories. The methods include a measure of distance between two such trees, and a measure of concordance between one such tree and a hierarchical classification tree of the unique categories. We demonstrate the intuition of our methods with some toy examples before presenting an analysis of Mycobacterium tuberculosis trees, in which we use our methods to quantify the differences between trees built from typing versus sequence data.

evolutionary biology

Genome-based transmission modeling separates imported tuberculosis from recent transmission within an immigrant population

BackgroundIn many countries tuberculosis incidence is low and largely shaped by immigrant populations from high-burden countries. This is the case in Norway, where more than 80 per cent of TB cases are found among immigrants from high-incidence countries. A variable latent period, low rates of evolution and structured social networks make separating import from within-border transmission a major conundrum to TB-control efforts in many low-incidence countries.\n\nMethodsClinical Mycobacterium tuberculosis isolates belonging to an unusually large genotype cluster associated with people born in the Horn of Africa, have been identified in Norway over the last two decades. We applied modeled transmission based on whole-genome sequence data to estimate infection times for individual patients. By contrasting these estimates with time of arrival in Norway, we estimate on a case-by-case basis whether patients were likely to have been infected before or after arrival.\n\nResultsIndependent import was responsible for the majority of cases, but we estimate that about a quarter of the patients had contracted TB in Norway.\n\nConclusionsThis study illuminates the transmission dynamics within an immigrant community. Our approach is broadly applicable to many settings where TB control programs can benefit from understanding when and where patients acquired tuberculosis.

epidemiology

Host population structure and treatment frequency maintain balancing selection on drug resistance

It is a truism that antimicrobial drugs select for resistance, but explaining pathogen- and population-specific variation in patterns of resistance remains an open problem. Like other common commensals, Streptococcus pneumoniae has demonstrated persistent coexistence of drug-sensitive and drug-resistant strains. Theoretically, this outcome is unlikely. We modeled the dynamics of competing strains of S. pneumoniae to investigate the impact of transmission dynamics and treatment-induced selective pressures on the probability of stable coexistence. We find that the outcome of competition is extremely sensitive to structure in the host population, although coexistence can arise from age-assortative transmission models with age-varying rates of antibiotic use. Moreover, we find that the selective pressure from antibiotics arises not so much from the rate of antibiotic use per se but from the frequency of treatment: frequent antibiotic therapy disproportionately impacts the fitness of sensitive strains. This same phenomenon explains why serotypes with longer durations of carriage tend to be more resistant. These dynamics may apply to other potentially pathogenic, microbial commensals and highlight how population structure, which is often omitted from models, can have a large impact.

evolutionary biology

Gene-centric constraint of metabolic models

MotivationA number of approaches have been introduced in recent years allowing gene expression data to be integrated into the standard flux Balance Analysis (FBA) technique. This additional information permits greater accuracy in the prediction of intracellular fluxes, even when knowledge of the growth medium and biomass composition is incomplete, and allows exploration of organisms metabolism under wide-ranging conditions. However, existing techniques still focus on the reaction as the fundamental unit of their modelling. This carries the advantages of incorporating expression measurements, but discounts the fact that genes (and their associated proteins) may be involved in the catalysis of multiple reactions through the formation of alternative protein complexes.\n\nResultsWe demonstrate an approach focusing not on reactions or genes as the fundamental unit, but on the Gene Complex (GC), a set of genes that is sufficient to catalyse a given reaction. We define expression-based limits in such a way that proteins cannot do double duty: no single molecule is permitted to contribute to the catalysis of more than one reaction at a time. Using experimentally determined RNA expression and intracellular fluxes, we validate this novel and more conceptually sound approach.\n\nAvailability and ImplementationAn implementation of the GC-FO_SCPLOWLUXC_SCPLOW algorithm is available as part of the Pyabolism python module. https://github.com/nickfyson/pyabolism\n\nContactnickfyson@gmail.com

biochemistry

Phylodynamics without trees: estimating R0 directly from pathogen sequences

We develop a new tree-free phylodynamic method to estimate the reproduction number (R0) of a pathogen from large numbers of sequences of a pathogen. It is based on the convergence of the cherry-to-tip ratio (CTR) to a constant depending on R0 in supercritical branching trees. It is a tree-free method because tree reconstruction is not required: the number of cherries and the CTR is estimated directly from the sequences using the new computational method Cherries Without Tree (CWT). With simulations, we compare CWT to other methods currently in use. We use the new inference method to estimate R0 from simulated sequences and discuss its accuracy. We explore the potential bias arising from sub-sampling.

epidemiology