Search bioRxivSearch

bioRxiv · 10.1101/069195

New method to reconstruct phylogenetic and transmission trees with sequence data from infectious disease outbreaks

Abstract

Whole-genome sequencing (WGS) of pathogens from host samples becomes more and more routine during infectious disease outbreaks. These data provide information on possible transmission events which can be used for further epidemiologic analyses, such as identification of risk factors for infectivity and transmission. However, the relationship between transmission events and WGS data is obscured by uncertainty arising from four largely unobserved processes: transmission, case observation, within-host pathogen dynamics and mutation. To properly resolve transmission events, these processes need to be taken into account. Recent years have seen much progress in theory and method development, but applications are tailored to specific datasets with matching model assumptions and code, or otherwise make simplifying assumptions that break up the dependency between the four processes. To obtain a method with wider applicability, we have developed a novel approach to reconstruct transmission trees with WGS data. Our approach combines elementary models for transmission, case observation, within-host pathogen dynamics, and mutation. We use Bayesian inference with MCMC for which we have designed novel proposal steps to efficiently traverse the posterior distribution, taking account of all unobserved processes at once. This allows for efficient sampling of transmission trees from the posterior distribution, and robust estimation of consensus transmission trees. We implemented the proposed method in a new R package phybreak. The method performs well in tests of both new and published simulated data. We apply the model to to five datasets on densely sampled infectious disease outbreaks, covering a wide range of epidemiological settings. Using only sampling times and sequences as data, our analyses confirmed the original results or improved on them: the more realistic infection times place more confidence in the inferred transmission trees.\n\nAuthor SummaryIt is becoming easier and cheaper to obtain whole genome sequences of pathogen samples during outbreaks of infectious diseases. If all hosts during an outbreak are sampled, and these samples are sequenced, the small differences between the sequences (single nucleotide polymorphisms, SNPs) give information on the transmission tree, i.e. who infected whom, and when. However, correctly inferring this tree is not straightforward, because SNPs arise from unobserved processes including infection events, as well as pathogen growth and mutation within the hosts. Several methods have been developed in recent years, but none so generic and easily accessible that it can easily be applied to new settings and datasets. We have developed a new model and method to infer transmission trees without putting prior limiting constraints on the order of unobserved events. The method is easily accessible in an R package implementation. We show that the method performs well on new and previously published simulated data. We illustrate applicability to a wide range of infectious diseases and settings by analysing five published datasets on densely sampled infectious disease outbreaks, confirming or improving the original results.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Don Klinkenberg, Jantien Backer, Xavier Didelot, Caroline Colijn, Jacco Wallinga. 2016-08-12. New method to reconstruct phylogenetic and transmission trees with sequence data from infectious disease outbreaks. https://doi.org/10.1101/069195

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A risk stratification approach for improved interpretation of diagnostic accuracy statistics

Diagnostic accuracy statistics, including predictive values, risk-differences, Youdens index and Area Under the Curve (AUC), assess the promise of novel biomarkers proposed as diagnostic tests. We reinterpret these statistics in light of risk-stratification (how well a biomarker separates those at higher risk from those at lower risk) to better understand their implications for public-health programs. We introduce an intuitively simple statistic, Mean Risk Stratification (MRS): the average change in risk (pre-test vs. post-test) revealed for tested individuals. High MRS implies better risk separation achieved by testing. MRS demonstrates that conventional predictive values can mislead because they do not account for disease prevalence or test-positivity rates. Little risk-stratification is possible for rare diseases, demonstrating a \"high-bar\" to justify population-based screening. Importantly, we demonstrate that the risk-difference, Youdens index, and AUC measure only multiplicative relative gains in risk-stratification: AUC=0.6 achieves only 20% of maximum risk-stratification (AUC=0.9 achieves 80%). However, large relative gains in risk-stratification might not imply large absolute gains if disease is rare or if the test is rarely positive. We illustrate MRS by our experience comparing the performance of cervical cancer screening tests in China vs. the USA. The test with the worst AUC=0.72 in China (visual inspection with ascetic acid) provides twice the risk-stratification of the test with best AUC=0.83 in the USA (human papillomavirus and Pap cotesting) because China has three times more cervical precancer/cancer. MRS could be routinely calculated to better understand the clinical/public-health implications of standard diagnostic accuracy statistics.

Epidemiology

Collider Scope: How selection bias can induce spurious associations

Large-scale cross-sectional and cohort studies have transformed our understanding of the genetic and environmental determinants of health outcomes. However, the representativeness of these samples may be limited - either through selection into studies, or by attrition from studies over time. Here we explore the potential impact of this selection bias on results obtained from these studies, from the perspective that this amounts to conditioning on a collider (i.e., a form of collider bias). While it is acknowledged that selection bias will have a strong effect on representativeness and prevalence estimates, it is often assumed that it should not have a strong impact on estimates of associations. We argue that because selection can induce collider bias (which occurs when two variables independently influence a third variable, and that third variable is conditioned upon), selection can lead to substantially biased estimates of associations. In particular, selection related to phenotypes can bias associations with genetic variants associated with those phenotypes. In simulations, we show that even modest influences on selection into, or attrition from, a study can generate biased and potentially misleading estimates of both phenotypic and genotypic associations. Our results highlight the value of knowing which population your study sample is representative of. If the factors influencing selection and attrition are known, they can be adjusted for. For example, having DNA available on most participants in a birth cohort study offers the possibility of investigating the extent to which polygenic scores predict subsequent participation, which in turn would enable sensitivity analyses of the extent to which bias might distort estimates.\n\nKey MessagesSelection bias (including selective attrition) may limit the representativeness of large-scale cross-sectional and cohort studies.\n\nThis selection bias may induce collider bias (which occurs when two variables independently influence a third variable, and that variable is conditioned upon).\n\nThis may lead to substantially biased estimates of associations, including of genetic associations, even when selection / attrition is relatively modest.

Epidemiology

A comparative analysis of Chikungunya and Zika transmission

The recent global dissemination of Chikungunya and Zika has fostered public health concern worldwide. To better understand the drivers of transmission of these two arboviral diseases, we propose a joint analysis of Chikungunya and Zika epidemics in the same territories, taking into account the common epidemiological features of the epidemics: transmitted by the same vector, in the same environments, and observed by the same surveillance systems. We analyse eighteen outbreaks in French Polynesia and the French West Indies using a hierarchical time-dependent SIR model accounting for the effect of virus, location and weather on transmission, and based on a disease specific serial interval. We show that Chikungunya and Zika have similar transmission potential in the same territories (transmissibility ratio between Zika and Chikungunya of 1.04 [95% credible interval: 0.97; 1.13]), but that detection and reporting rates were different (around 19% for Zika and 40% for Chikungunya). Temperature variations between 22{degrees}C and 29{degrees}C did not alter transmission, but increased precipitation showed a dual effect, first reducing transmission after a two-week delay, then increasing it around five weeks later. The present study provides valuable information for risk assessment and introduces a modelling framework for the comparative analysis of arboviral infections that can be extended to other viruses and territories.

Epidemiology