Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Epidemiology of Cancers in Zambia: A Significant Variation in Cancer Incidence and Prevalence across the Nation

BackgroundCancers are one of the leading causes of death worldwide. More than two thirds of deaths due to cancers occur in low- and middle-income countries whereZambia belongs. This study therefore sought to assess the epidemiology of cancers in Zambia.\n\nMethodsWe conducted a retrospective observational study nested on Zambia National Cancer Registry (ZNCR) histopathological and clinical data from 2007 to 2014. Zambia Central Statistics Office (CSO)demographic datawere used to calculate prevalence and incidence rates of cancers. Age-adjusted rates and case fatality rates were estimated using standard methods. We used a Poisson Approximation for calculating 95% confidence intervals (CI).\n\nResultsThe top seven most cancer prevalent districts in Zambia have been Luangwa, Kabwe, Lusaka, Monze, Mongu, Katete and Chipata. Cervical cancer, prostate cancer, breast cancer and Kaposis sarcoma were the top four most prevalent cancers as well as major causes of cancer related deaths in Zambia. Standardised Incidence Rates and 95% CI for the top four cancers were: cervix uteri (186.3; CI = 181.77 - 190.83), prostate (60.03; CI = 57.03 - 63.03), breast (38.08; CI = 36.0 - 40.16) and Kaposis sarcoma (26.18; CI = 25.14 - 27.22).CFR were: Leukaemia (38.1%); pancreatic cancer (36.3%); lung cancer (33.3%); and brain, nervous system (30.2%). Cancers were associated with HIV with p-value of 0.000 and Pearson correlation coefficient of 0.818.\n\nConclusionsThe widespread distribution of cancers with high prevalence in the southern zone has been perpetrated by lifestyle and sexual culture as well as geography. Intensifying cancer screening and early detection countrywide as well as changing the lifestyle and sexual culture would greatly help in the reduction of cancer cases in Zambia.

epidemiology

Bayesian learning ecosystem dynamics with delayed dependencies from incomplete multiple source data : an application to plant epidemiology

Ecosystem dynamics forecasting is central to major problems in ecology, society, and economy. The existing models serve as decision tools but their parameters valitity are usually not confronted to real data in a formalized approach. Dynamics bayesian network inference is promissing but limited when dealing with incomplete multiple source time series with delayed time dependencies. We propose here a temporal bayesian network with time delay and aproximate inference algorithm, to learn altogether cryptic ecosystem variables, missing data, and model parameters. The novelty in the approach is that it combines simulation-based and likelihood-based aproximate bayesian inference. The advantage of simulation based is that it allows to sample hidden processes. The advantage of likelihood based is that it provides a summary statistics that is really representing the model we are interested in. The ecosystem variables and the missing data are simulated from indicator variables using the probabilistic indicator-ecosystem model. The likelihood is estimated by averaging the probability of observed-simulated data over simulations, the parameter space is sampled with Metropolis Hasting algorithm. Another innovative proposition is to parametrize the network structure in order to learn model structure within a space provided by prior distribution. We apply to plant epidemiology.

epidemiology

Spatial epidemiology of networked metapopulation: An overview

An emerging disease is one infectious epidemic caused by a newly transmissible pathogen, which has either appeared for the first time or already existed in human populations, having the capacity to increase rapidly in incidence as well as geographic range. Adapting to human immune system, emerging diseases may trigger large-scale pandemic spreading, such as the transnational spreading of SARS, the global outbreak of A(H1N1), and the recent potential invasion of avian influenza A(H7N9). To study the dynamics mediating the transmission of emerging diseases, spatial epidemiology of networked metapopulation provides a valuable modeling framework, which takes spatially distributed factors into consideration. This review elaborates the latest progresses on the spatial metapopulation dynamics, discusses empirical and theoretical findings that verify the validity of networked metapopulations, and the application in evaluating the effectiveness of disease intervention strategies as well.

Biophysics

Epidemiological and evolutionary analysis of the 2014 Ebola virus outbreak

The 2014 epidemic of the Ebola virus is governed by a genetically diverse viral population. In the early Sierra Leone outbreak, a recent study has identified new mutations that generate genetically distinct sequence clades [1]. Here we find evidence that major Sierra Leone clades have systematic differences in growth rate and reproduction number. If this growth heterogeneity remains stable, it will generate major shifts in clade frequencies and influence the overall epidemic dynamics on time scales within the current outbreak. Our method is based on simple summary statistics of clade growth, which can be inferred from genealogical trees with an underlying clade-specific birth-death model of the infection dynamics. This method can be used to perform realtime tracking of an evolving epidemic and identify emerging clades of epidemiological or evolutionary significance.

Evolutionary Biology

Genomic epidemiology of the current wave of artemisinin resistant malaria

Artemisinin resistant Plasmodium falciparum is advancing across Southeast Asia in a soft selective sweep involving at least 20 independent kelch13 mutations. In a large global survey, we find that kelch13 mutations which cause resistance in Southeast Asia are present at low frequency in Africa. We show that African kelch13 mutations have originated locally, and that kelch13 shows a normal variation pattern relative to other genes in Africa, whereas in Southeast Asia there is a great excess of non-synonymous mutations, many of which cause radical amino-acid changes. Thus, kelch13 is not currently undergoing strong selection in Africa, despite a deep reservoir of standing variation that could potentially allow resistance to emerge rapidly. The practical implications are that public health surveillance for artemisinin resistance should not rely on kelch13 data alone, and interventions to prevent resistance must account for local evolutionary conditions, shown by genomic epidemiology to differ greatly between geographical regions.

Evolutionary Biology

DBSIM: A Platform of Simulation Resources for Genetic Epidemiology Studies

Computer simulations are routinely conducted to evaluate new statistical methods, to compare the properties among different methods, and to mimic the real data in genetic epidemiology studies. Conducting simulation studies can become a complicated task as several challenges can occur, such as the selection of an appropriate simulation tool and the specification of parameters in the simulation model. Although abundant simulated data have been generated for human genetic research, currently there is no public database designed specifically as a repository for these simulated data. With the lack of such database, for similar studies, similar simulations may have been repeated, which resulted in redundant works. We created an online platform, DBSIM, for simulation data sharing and discussion of simulation techniques for human genetic studies. DBSIM has a database containing simulation scripts, simulated data, and documentations from published manuscripts, as well as a discussion forum, which provides a platform for discussion of the simulated data and exchanging simulation ideas. DBSIM will be useful in three aspects. Moreover, summary statistics such as the simulation tools that are most commonly used and datasets that are most frequently downloaded are provided. The statistics will be very informative for researchers to choose an appropriate simulation tool or select a common dataset for method comparisons. DBSIM can be accessed at http://dbsim.nhri.org.tw.

Bioinformatics

16S rRNA amplicon sequencing for epidemiological surveys of bacteria in wildlife: the importance of cleaning post-sequencing data before estimating positivity, prevalence and co-infection

ImportanceSeveral recent public health crises have shown that the surveillance of zoonotic agents in wildlife is important to prevent pandemic risks. Rodents are intermediate hosts for numerous zoonotic bacteria. High-throughput sequencing (HTS) technologies are very useful for the detection and surveillance of zoonotic bacteria, but rigorous experimental processes are required for the use of these cheap and effective tools in such epidemiological contexts. In particular, HTS introduces biases into the raw dataset that might lead to incorrect interpretations. We describe here a procedure for cleaning data before estimating reliable biological parameters, such as bacterial positivity, prevalence and coinfection, by 16S rRNA amplicon sequencing on the MiSeq platform. This procedure, applied to 711 commensal rodents collected from 24 villages in Senegal, Africa, detected several emerging bacterial genera, some in high prevalence, while never before reported for West Africa. This study constitutes a step towards the use of HTS to improve our understanding of the risk of zoonotic disease transmission posed by wildlife, by providing a new strategy for the use of HTS platforms to monitor both bacterial diversity and infection dynamics in wildlife. In the future, this approach could be adapted for the monitoring of other microbes such as protists, fungi, and even viruses.\n\nSummaryHuman impact on natural habitats is increasing the complexity of human-wildlife interfaces and leading to the emergence of infectious diseases worldwide. Highly successful synanthropic wildlife species, such as rodents, will undoubtedly play an increasingly important role in transmitting zoonotic diseases. We investigated the potential of recent developments in 16S rRNA amplicon sequencing to facilitate the multiplexing of large numbers of samples, to improve our understanding of the risk of zoonotic disease transmission posed by urban rodents in West Africa. In addition to listing pathogenic bacteria in wild populations, as in other high-throughput sequencing (HTS) studies, our approach can estimate essential parameters for studies of zoonotic risk, such as prevalence and patterns of coinfection within individual hosts. However, the estimation of these parameters requires cleaning of the raw data to eliminate the biases generated by HTS methods. We present here an extensive review of these biases and of their consequences, and we propose a trimming strategy for managing them and cleaning the dataset. We also analyzed 711 commensal rodents collected from 24 villages in Senegal, including 208 Mus musculus domesticus, 189 Rattus rattus, 93 Mastomys natalensis and 221 Mastomys erythroleucus. Seven major genera of pathogenic bacteria were detected: Borrelia, Bartonella, Mycoplasma, Ehrlichia, Rickettsia, Streptobacillus and Orientia. The last five of these genera have never before been detected in West African rodents. Bacterial prevalence ranged from 0% to 90%, depending on the bacterial taxon, rodent species and site considered, and a mean of 26% of rodents displayed coinfection. The 16S rRNA amplicon sequencing strategy presented here has the advantage over other molecular surveillance tools of dealing with a large spectrum of bacterial pathogens without requiring assumptions about their presence in the samples. This approach is, thus, particularly suitable for continuous pathogen surveillance in the framework of disease monitoring programs

Microbiology

Population genomics of bank vole populations reveals associations between immune related genes and the epidemiology of Puumala hantavirus in Sweden

Infectious pathogens are major selective forces acting on individuals. The recent advent of high-throughput sequencing technologies now enables to investigate the genetic bases of resistance/susceptibility to infections in non-model organisms. From an evolutionary perspective, the analysis of the genetic diversity observed at these genes in natural populations provides insight into the mechanisms maintaining polymorphism and their epidemiological consequences. We explored these questions in the context of the interactions between Puumala hantavirus (PUUV) and its reservoir host, the bank vole Myodes glareolus. Despite the continuous spatial distribution of M. glareolus in Europe, PUUV distribution is strongly heterogeneous. Different defence strategies might have evolved in bank voles as a result of co-adaptation with PUUV, which may in turn reinforce spatial heterogeneity in PUUV distribution. We performed a genome scan study of six bank vole populations sampled along a North/South transect in Sweden, including PUUV endemic and non-endemic areas. We combined candidate gene analyses (Tlr4, Tlr7, Mx2 genes) and high throughput sequencing of RAD (Restriction-site Associated DNA) markers. We found evidence for outlier loci showing high levels of genetic differentiation. Ten outliers among the 52 that matched to mouse protein-coding genes corresponded to immune related genes and were detected using ecological associations with variations in PUUV prevalence. One third of the enriched pathways concerned immune processes, including platelet activation and TLR pathway. In the future, functional experimentations should enable to confirm the role of these these immune related genes with regard to the interactions between M. glareolus and PUUV.

evolutionary biology

The influence of landscape and environmental factors on ranavirus epidemiology in amphibian assemblages

AimTo quantify the influence of a suite of landscape, abiotic, biotic, and host-level variables on ranavirus disease dynamics in amphibian assemblages at two biological levels (site and host-level).\n\nLocationWetlands within the East Bay region of California, USA.\n\nMethodsWe used competing models, multimodel inference, and variance partitioning to examine the influence of 16 landscape and environmental factors on patterns in site-level ranavirus presence and host-level ranavirus infection in 76 wetlands and 1,377 amphibian hosts representing five species.\n\nResultsThe landscape factor explained more variation than any other factors in site-level ranavirus presence, but biotic and host-level factors explained more variation in host-level ranavirus infection. At both the site- and host-level, the probability of ranavirus presence correlated negatively with distance to nearest ranavirus-positive wetland. At the site-level, ranavirus presence was associated positively with taxonomic richness. However, infection prevalence within the amphibian population correlated negatively with vertebrate richness. Finally, amphibian host species differed in their likelihood of ranavirus infection: American Bullfrogs had the weakest association with infection while Western Toads had the strongest. After accounting for host species effects, hosts with greater snout-vent length had a lower probability of infection.\n\nMain conclusionsStrong spatial influences at both biological levels suggest that mobile taxa (e.g., adult amphibians, birds, reptiles) may facilitate the movement of ranavirus among hosts and across the landscape. Higher taxonomic richness at sites may provide more opportunities for colonization or the presence of reservoir hosts that may influence ranavirus presence. Higher host richness correlating with higher ranavirus infection is suggestive of a dilution effect that has been observed for other amphibian disease systems and warrants further investigation. Our study demonstrates that an array of landscape, environmental, and host-level factors were associated with ranavirus epidemiology and illustrates that their importance vary with biological level.

ecology

Microbiomes of North American Triatominae: the grounds for Chagas disease epidemiology

AbstarctInsect microbiomes influence many fundamental host traits, including functions of practical significance such as their capacity as vectors to transmit parasites and pathogens. The knowledge on the diversity and development of the gut microbiomes in various blood feeding insects is thus crucial not only for theoretical purposes, but also for the development of better disease control strategies. In Triatominae (Heteroptera: Reduviidae), the blood feeding vectors of Chagas disease in South America and parts of North America, the investigation of the microbiomes is in its infancy. The few studies done on microbiomes of South American Triatominae species indicate a relatively low taxonomic diversity and a high host specificity. We designed a comparative survey to serve several purposes: I) to obtain a better insight into the overall microbiome diversity in different species, II) to check the long term stability of the interspecific differences, III) to describe the ontogenetic changes of the microbiome, and IV) to determine the potential correlation between microbiome composition and presence in the vector gut of Trypanosoma cruzi, the causative agent of Chagas disease. Using 16S amplicons of two abundant species from the southern US, and four laboratory reared colonies, we showed that the microbiome composition is determined by host species, rather than locality or environment. The OTUs (Operational Taxonomic Units) determination confirms a low microbiome diversity, with 12-17 main OTUs detected in wild populations of T. sanguisuga and T. protracta. Among the dominant bacterial taxa are Acinetobacter and Proteiniphilum but also the symbiotic bacterium Arsenophonus triatominarum, previously believed to only live intracellularly. The possibility of ontogenic microbiome changes was evaluated in all six developmental stages and faeces of the laboratory reared model Rhodnius prolixus. We detected considerable changes along the hosts ontogeny, including clear trends in the abundance variation of the three dominant bacteria, namely Enterococcus, Acinetobacter and Arsenophonus. Finally, we screened the samples for the presence of T. cruzi. Comparing the parasite presence with the microbiome composition, we assessed the possible significance of the latter in the epidemiology of the disease. Particularly, we found a trend towards more diverse microbiomes in T. cruzi positive T. protracta specimens.

microbiology

Epidemiology, Microbiology and Therapeutic Consequences of Chronic Osteomyelitis in Northern China: A Retrospective Analysis of 255 Patients

The study aimed to explore the epidemiology and clinical characteristics of chronic osteomyelitis observed in a northern China hospital. Clinical data of 255 patients with chronic osteomyelitis from January 2007 to January 2014 were collected and analyzed, including general information, disease data, treatment and follow-up data. Chronic osteomyelitis is more common in males and in the age group from 41-50 years of age. Common infection sites are the femur, tibiofibular, and hip joint. More g+ than g- bacterial infections were observed, with S. aureus the most commonly observed pathogenic organism. The positive detection rate from debridement bacterial culture is 75.6%. The detection rate when five samples are sent for bacterial culture is 90.6%, with pathogenic bacteria identified in 82.8% of cases. The two-stage debridement method (87.0%) has higher first curative rate than the one-stage debridement method (71.2%). To improve detection rate using bacterial culture, at least five samples are recommended. Treatment of chronic osteomyelitis with two-stage debridement, plus antibiotic-loaded polymethylmethacrylate (PMMA) beads provided good clinical results in this study and is therefore recommended.

microbiology

Mutators as drivers of adaptation in pathogenic bacteria and a risk factor for host jumps and vaccine escape: insights into evolutionary epidemiology of the global aquatic pathogen Streptococcus iniae

Pathogens continuously adapt to changing host environments where variation in their virulence and antigenicity is critical to their long-term evolutionary success. The emergence of novel variants is accelerated in microbial mutator strains (mutators) deficient in DNA repair genes, most often from mismatch repair and oxidised-guanine repair systems (MMR and OG respectively). Bacterial MMR/OG mutants are abundant in clinical samples and show increased adaptive potential in experimental infection models, yet the role of mutators in the epidemiology and evolution of infectious disease is not well understood. Here we investigated the role of mutation rate dynamics in the evolution of a broad host range pathogen, Streptococcus iniae, using a set of 80 strains isolated globally over 40 years. We have resolved phylogenetic relationships using non-recombinant core genome variants, measured in vivo mutation rates by fluctuation analysis, identified variation in major MMR/OG genes and their regulatory regions, and phenotyped the major traits determining virulence in streptococci. We found that both mutation rate and MMR/OG genotype are remarkably conserved within phylogenetic clades but significantly differ between major phylogenetic lineages. Further, variation in MMR/OG loci correlates with occurrence of atypical virulence-associated phenotypes, infection in atypical hosts (mammals), and atypical tissue of a vaccinated primary hosts (barramundi bone). These findings suggest that mutators are likely to facilitate adaptations preceding major diversification events, and may promote emergence of variation permitting colonisation of a novel host tissue, novel host taxa (host jumps), and immune-escape in the vaccinated host.

microbiology

Fast and flexible bacterial genomic epidemiology with PopPUNK

The routine use of genomics for disease surveillance provides the opportunity for high-resolution bacterial epidemiology.\n\nHowever, current whole-genome clustering and multi-locus typing approaches do not fully exploit core and accessory genomic variation, and cannot both automatically identify, and subsequently expand, clusters of significantly-similar isolates in large datasets and across species.\n\nHere we describe PopPUNK (Population Partitioning Using Nucleotide K-mers; https://poppunk.readthedocs.io/en/latest/). software implementing scalable and expandable annotation- and alignment-free methods for population analysis and clustering.\n\nVariable-length k-mer comparisons are used to distinguish isolates divergence in shared sequence and gene content, which we demonstrate to be accurate over multiple orders of magnitude using both simulated data and real datasets from ten taxonomically-widespread species. Connections between closely-related isolates of the same strain are robustly identified, despite variation in the discontinuous pairwise distance distributions that reflects species diverse evolutionary patterns. PopPUNK can process 103-104 genomes as single batch, with minimal memory use and runtimes up to 200-fold faster than existing methods. Clusters of strains remain consistent as new batches of genomes are added, which is achieved without needing to re-analyse all genomes de novo.\n\nThis facilitates real-time surveillance with stable cluster naming and allows for outbreak detection using hundreds of genomes in minutes. Interactive visualisation and online publication is streamlined through automatic output of results to multiple platforms.\n\nPopPUNK has been designed as a flexible platform that addresses important issues with currently used whole-genome clustering and typing methods, and has potential uses across bacterial genetics and public health research.

genomics

Epidemiological impact of hepatitis B vaccination in Monastir Tunisia (2000-17)

Background: In 2016, the first global health sector strategy on viral hepatitis was endorsed with the goal of eliminating viral hepatitis as a public health threat by 2030. In Tunisia, effective vaccines for hepatitis B (HBV) have been added to the expanded programme of immunization (EPI) since July 1995 for new borns. We expected to have a decreasing trend in the prevalence rate of reported HBV. Our study aimed to address the epidemiological profile of HBV, to assess trends by age and gender in Monastir governorate over a period of 18 years according to immunization status and to estimate the burden (years lived with disability YLDs) of this pathology. Methods: We performed a descriptive cross sectional study of declared HBV from January 1, 2000 to December 31, 2017 defined as having positive serologic markers for HBs Ag. All declared patients were residents of Monastirs Governorate. EPI included two periods, the first between 1995 and 2006 following a three-dose schedule (3, 4, 9 months). The second PI cohort after 2006 following a three-dose schedule (0, 2, 6 months). Results: During 18 years, 1526 cases of HBV were declared in Monastir with a mean of 85 cases per year. We estimated a mean of 1699 declared cases per year of HBV in Tunisia. CPR was 16.85/100,000 inh being the higher in age group of 20-39 years and in men .ASR was 15.99/100,000 inh, being 35.5 in men and 8.69 in women. During the study period, declared cases among presumed immunized (PI) person against HBV were 32(2.0%). Among PI cases, 29 were from the first period and 3 were from the second. We established a negative trend over 18 years of hepatitis B. The age group of 20 to 39 was the most common with a sharply decline. Presumed not immunized (PNI) HBV cases are decreasing by years with a prediction of 35 cases in 2024. Reported HBV contributed to 1.26 YLDs per 100,000 inh. The highest rate of YLDs occurred at the age 20-39 (2.73 YLDs per 100,000 inh). During 18 years, YLDs were 114.45 in Monastir with a mean of 2293.65 YLDs of HBV in Tunisia. Conclusion, this study showed a law prevalence rate and a decreasing trend of HBV during 18 years showing an efficacy of immunization and confirming that the universal hepatitis B vaccination in Tunisia has resulted in progress towards the prevention and control of hepatitis B infection. These findings should be demonstrated in other Tunisian regions with a standardized serological profile.

immunology

Tracking a Serial Killer: Integrating Phylogenetic Relationships, Epidemiology, and Geography for Two Invasive Meningococcal Disease Outbreaks

BackgroundWhile overall rates of meningococcal disease have been declining in the United States for the past several decades, New York City (NYC) has experienced two serogroup C meningococcal disease outbreaks in 2005-2006 and in 2010-2013. The outbreaks were centered within drug use and sexual networks, were difficult to control, and required vaccine campaigns.\n\nMethodsWhole Genome Sequencing (WGS) was used to analyze preserved meningococcal isolates collected before and during the two outbreaks. We integrated and analyzed epidemiologic, geographic, and genomic data to better understand transmission networks among patients. Betweenness centrality was used as a metric to understand the most important geographic nodes in the transmission networks. Comparative genomics was used to identify genes associated with the outbreaks.\n\nResultsNeisseria meningitidis serogroup C (ST11/ET-37) was responsible for both outbreaks with each outbreak having distinct phylogenetic clusters. WGS did identify some misclassifications of isolates that were more distant from the rest of the outbreak, as well as those that should have been included based on high genomic similarity. Genomes for the second outbreak were more similar than the first and no mutation was found to either be unique or specific to either outbreak lineage. Betweenness centrality as applied to transmission networks based on phylogenetic analysis demonstrated that the outbreaks were transmitted within focal communities in NYC with few transmission events to other locations.\n\nConclusionsNeisseria meningitidis is an ever changing pathogen and comparative genomic analyses can help elucidate how it spreads geographically to facilitate targeted interventions to interrupt transmission.

genomics

Genotypic clustering does not imply recent tuberculosis transmission in a high prevalence setting: A genomic epidemiology study in Lima, Peru

BackgroundWhole genome sequencing (WGS) can elucidate Mycobacterium tuberculosis (Mtb) transmission patterns but more data is needed to guide its use in high-burden settings. In a household-based transmissibility study of 4,000 TB patients in Lima, Peru, we identified a large MIRU-VNTR Mtb cluster with a range of resistance phenotypes and studied host and bacterial factors contributing to its spread.\n\nMethodsWGS was performed on 61 of 148 isolates in the cluster. We compared transmission link inference using epidemiological or genomic data with and without the inclusion of controversial variants, and estimated the dates of emergence of the cluster and antimicrobial drug resistance acquisition events by generating a time-calibrated phylogeny. We validated our findings in genomic data from an outbreak of 325 TB cases in London. Using a larger set of 12,032 public Mtb genomes, we determined bacterial factors characterizing this cluster and under positive selection in other Mtb lineages.\n\nFindingsFour isolates were distantly related and the remaining 57 isolates diverged ca. 1968 (95% HPD: 1945-1985). Isoniazid resistance arose once, whereas rifampicin resistance emerged subsequently at least three times. Amplification of other drug resistance occurred as recently as within the last year of sampling. High quality PE/PPE variants and indels added information for transmission inference. We identified five cluster-defining SNPs, including esxV S23L to be potentially contributing to transmissibility.\n\nInterpretationClusters defined by MIRU-VNTR typing, could be circulating for decades in a high-burden setting. WGS allows for an improved understanding of transmission, as well as bacterial resistance and fitness factors.\n\nFundingThe study was funded by the National Institutes of Health (Peru Epi study U19-AI076217 and K01-ES026835 to MRF). The funding sources had no role in any aspect of the study, manuscript or decision to submit it for publication.\n\nResearch in contextO_ST_ABSEvidence before this studyC_ST_ABSUse of whole genome sequencing (WGS) to study tuberculosis (TB) transmission has proven to have higher resolution that traditional typing methods in low-burden settings. The implications of its use in high-burden settings are not well understood.\n\nAdded value of this studyUsing WGS, we found that TB clusters defined by traditional typing methods may be circulating for several decades. Genomic regions typically excluded from WGS analysis contain large amount of genetic variation that may affect interpretation of transmission events. We also identified five bacterial mutations that may contribute to transmission fitness.\n\nImplications of all the available evidenceAdded value of WGS for understanding TB transmission may be even higher in high-burden vs. low-burden settings. Methods integrating variants found in polymorphic sites and insertions and deletions are likely to have higher resolution. Several host and bacterial factors may be responsible for higher transmissibility that can be targets of intervention to interrupt TB transmission in communities.

microbiology

Probabilistic forecasting in infectious disease epidemiology: The thirteenth Armitage lecture

Routine surveillance of notifiable infectious diseases gives rise to daily or weekly counts of reported cases stratified by region and age group. From a public health perspective, forecasts of infectious disease spread are of central importance. We argue that such forecasts need to properly incorporate the attached uncertainty, so should be probabilistic in nature. However, forecasts also need to take into account temporal dependencies inherent to communicable diseases, spatial dynamics through human travel, and social contact patterns between age groups. We describe a multivariate time series model for weekly surveillance counts on norovirus gastroenteritis from the 12 city districts of Berlin, in six age groups, from week 2011/27 to week 2015/26. The following year (2015/27 to 2016/26) is used to assess the quality of the predictions. Probabilistic forecasts of the total number of cases can be derived through Monte Carlo simulation, but first and second moments are also available analytically. Final size forecasts as well as multivariate forecasts of the total number of cases by age group, by district, and by week are compared across different models of varying complexity. This leads to a more general discussion of issues regarding modelling, prediction and evaluation of public health surveillance data.

epidemiology

Multiple introductions of Zika virus into the United States revealed through genomic epidemiology

Zika virus (ZIKV) is causing an unprecedented epidemic linked to severe congenital syndromes1,2. In July 2016, mosquito-borne ZIKV transmission was first reported in the continental United States and since then, hundreds of locally-acquired infections have been reported in Florida3. To gain insights into the timing, source, and likely route(s) of introduction of ZIKV into the continental United States, we tracked the virus from its first detection in Miami, Florida by direct sequencing of ZIKV genomes from infected patients and Aedes aegypti mosquitoes. We show that at least four distinct ZIKV introductions contributed to the outbreak in Florida and that local transmission likely started in the spring of 2016 - several months before its initial detection. By analyzing surveillance and genetic data, we discovered that ZIKV moved among transmission zones in Miami. Our analyses show that most introductions are phylogenetically linked to the Caribbean, a finding corroborated by the high incidence rates and traffic volumes from the region into the Miami area. By comparing mosquito abundance and travel flows, we describe the areas of southern Florida that are especially vulnerable to ZIKV introductions. Our study provides a deeper understanding of how ZIKV initiates and sustains transmission in new regions.

epidemiology