Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The Kendrick Modelling Platform: Language Abstractions and Tools for Epidemiology

BackgroundMathematical and computational models are widely used for examining transmission, pathogenicity and propagation of infectious diseases. Software implementation of such models and their accompanied modelling tools do not adhere to well established software engineering principles. These principles include both modularity and clear separation of concerns that can promote reproducibility and reusability by other researchers. On the contrary the software written for epidemiology is monolithic, highly coupled and severely heterogeneous. This reality ultimately makes these computational models hard to study and to reuse, both because of the programming competence required and because of the incompatibility of the different approaches involved. Our goal with Kendrick is to simplify the creation of epidemiological models through a unified Domain-Specific Language for epidemiology that can support a variety of modelling and simulation approaches classically used in the field. This goal can be achieved by promoting reproducibility and reuse with modular modelling abstractions.\n\nResultsWe show through several examples how our modular abstractions and tools can reproduce uniformly complex mathematical and computational models of epidemics, despite being simulated by different methods. This is achieved without requiring sophisticated programming skills from the part of the user. We then successfully validate each kind of simulation through statistical analysis between the time series generated and the known theoretical expectations.\n\nConclusionsKendrick is one of the few DSLs for epidemiology that does not burden its users with implementation details or expecting sophisticated programming skills. It is also currently the only language for epidemiology that supports modularity through clear separation of concerns that promote reproducibility and reuse. Kendricks wider adoption and further development from the epidemiological community could boost research productivity in epidemiology by allowing researchers to easily reproduce and reuse each others software models and simulations. The tool can also be used by people who are not necessarily epidemiology modelers.

epidemiology

Transfer entropy as a tool for inferring causality from observational studies in epidemiology

Recently Wieners causality theorem, which states that one variable could be regarded as the cause of another if the ability to predict the future of the second variable is enhanced by implementing information about the preceding values of the first variable, was linked to information theory through the development of a novel metric called transfer entropy. Intuitively, transfer entropy can be conceptualized as a model-free measure of directed information flow from one variable to another. In contrast, directionality of information flow is not reflected in traditional measures of association which are completely symmetric by design. Although information theoretic approaches have been applied before in epidemiology, their value for inferring causality from observational studies is still unknown. Therefore, in the present study we use a set of simulation experiments, reflecting the most classical and widely used epidemiological observational study design, to validate the application of transfer entropy in epidemiological research. Moreover, we illustrate the practical applicability of this information theoretic approach to real-world epidemiological data by demonstrating that transfer entropy is able to extract the correct direction of information flow from longitudinal data concerning two well-known associations, i.e. that between smoking and lung cancer and that between obesity and diabetes risk. In conclusion, our results provide proof-of-concept that the recently developed transfer entropy method could be a welcome addition to the epidemiological armamentarium, especially to dissect those situations in which there is a well-described association between two variables but no clear-cut inclination as to the directionality of the association.

epidemiology

Proof of concept for quantitative urine NMR metabolomics pipeline for large-scale epidemiology and genetics

BackgroundQuantitative molecular data from urine are rare in epidemiology and genetics. NMR spectroscopy could provide these data in high-throughput, and it has already been applied in epidemiological settings to analyse urine samples. However, quantitative protocols for large-scale applications are not available.\n\nMethodsWe describe in detail how to prepare urine samples and perform NMR experiments to obtain quantitative metabolic information. Semi-automated quantitative lineshape fitting analyses were set up for 43 metabolites and applied to data from various analytical test samples and from 1,004 individuals from a population-based epidemiological cohort. Novel analyses on how urine metabolites associate with quantitative serum NMR metabolomics data (61 metabolic measures; n=995) were performed. In addition, confirmatory genome-wide analyses of urine metabolites were conducted (n=578). The fully automated quantitative regression-based spectral analysis is demonstrated for creatinine and glucose (n= 4,548).\n\nResultsIntra-assay metabolite variations were mostly <5% indicating high robustness and accuracy of the urine NMR spectroscopy methodology per se. Intra-individual metabolite variations were large, ranging from 6% to 194%. However, population-based inter-individual metabolite variations were even larger (from 14% to 1655%), providing a sound base for epidemiological applications. Metabolic associations between urine and serum were found clearly weaker than those within serum and within urine, indicating that urinary metabolomics data provide independent metabolic information. Two previous genome-wide hits for formate and 2-hydroxyisobutyrate were replicated at genome-wide significance.\n\nConclusionsQuantitative urine metabolomics data suggest broad novelty for systems epidemiology. A roadmap for an open access methodology is provided.

epidemiology

Real-time genomic and epidemiological investigation of a multi-institution outbreak of KPC-producing Enterobacteriaceae: a translational study

BackgroundUntil recently, KPC-producing Enterobacteriaceae were rarely identified in Australia. Following an increase in the number of incident cases across the state of Victoria, we undertook a real-time combined genomic and epidemiological investigation. The scope of this study included identifying risk factors and routes of transmission, and investigating the utility of genomics to enhance traditional field epidemiology for informing management of established widespread outbreaks.\n\nMethods and FindingsAll KPC-producing Enterobacteriaceae isolates referred to the state reference laboratory from 2012 onwards were included. Whole-genome sequencing (WGS) was performed in parallel with a detailed descriptive epidemiological investigation of each case, using Illumina sequencing on each isolate. This was complemented with PacBio long-read sequencing on selected isolates to establish high-quality reference sequences and interrogate characteristics of KPC-encoding plasmids. Initial investigations indicated the outbreak was widespread, with 86 KPC-producing Enterobacteriaceae isolates (K. pneumoniae 92%) identified from 35 different locations across metropolitan and rural Victoria between 2012-2015. Initial combined analyses of the epidemiological and genomic data resolved the outbreak into distinct nosocomial transmission networks, and identified healthcare facilities at the epicentre of KPC transmission. New cases were assigned to transmission networks in real-time, allowing focussed infection control efforts. PacBio sequencing confirmed a secondary transmission network arising from inter-species plasmid transmission. Insights from Bayesian transmission inference and analyses of within-host diversity informed the development of state-wide public health and infection control guidelines, including interventions such as an intensive approach to screening contacts following new case detection to minimise unrecognised colonisation.\n\nConclusionsA real-time combined epidemiological and genomic investigation proved critical to identifying and defining multiple transmission networks of KPC Enterobacteriaceae, while data from either investigation alone were inconclusive. The investigation was fundamental to informing infection control measures in real-time and the development of state-wide public health guidelines on carbapenemase producing Enterobacteriaceae management.

genomics

The Epidemiology and Transmissibility of Zika Virus in Girardot and San Andres Island, Colombia

BackgroundZika virus (ZIKV) is an arbovirus in the same genus as dengue virus and yellow fever virus. ZIKV transmission was first detected in Colombia in September 2015. The virus has spread rapidly across the country in areas infested with the vector Aedes aegypti. As of March 2016, Colombia has reported over 50,000 cases of Zika virus disease (ZVD).\n\nMethodsWe analyzed surveillance data of ZVD cases reported to the local health authorities of San Andres, Colombia, and Girardot, Colombia, between September 2015 and January 2016. Standardized case definitions used in both areas were determined by the Ministry of Health and Colombian National Institute of Health at the beginning of the ZIKV epidemic. ZVD was laboratory-confirmed by a finding of Zika virus RNA in the serum of acute cases. We report epidemiological summaries of the two outbreaks. We also use daily incidence data to estimate the basic reproductive number R0 in each population.\n\nFindingsWe identified 928 and 1,936 laboratory or clinically confirmed cases in San Andres and Girardot, respectively. The overall attack rate for reported ZVD detected by healthcare local surveillance was 12{middle dot}13 cases per 1,000 residents of San Andres and 18{middle dot}43 cases per 1,000 residents of Girardot. Attack rates were significantly higher in females in both municipalities. Cases occurred in all age groups but the most affected group was 20 to 49 year olds. The estimated R0 for the Zika outbreak in San Andres was 1{middle dot}41 (95% CI 1{middle dot}15 to 1{middle dot}74), and in Girardot was 4{middle dot}61 (95% CI 4{middle dot}11 to 5{middle dot}16).\n\nInterpretationTransmission of ZIKV is ongoing and spreading throughout the Americas rapidly. The observed rapid spread is supported by the relatively high basic reproductive numbers calculated from these two outbreaks in Colombia.\n\nFundingThis work was supported by National Institutes of Health (NIH) U54 GM111274, NIH R37 AI032042 and the Colombian Department of Science and Technology (Fulbright-Colciencias scholarship to D.P.R). The funding source had no role in the preparation of this manuscript or in the decision to publish this study.\n\nResearch in ContextO_ST_ABSEvidence before this studyC_ST_ABSThe ongoing outbreak of Zika virus disease in the Americas is the largest ever recorded. Since its first detection in April 2015 in Brazil, around 500,000 cases have been estimated, and the virus is spreading rapidly in the Americas region. There are many unanswered questions about the transmissibility and pathogenicity of the virus. Limited data are available from recent outbreaks occurring in islands in the Pacific, and little epidemiological data is available on the current outbreak.\n\nWe searched PubMed on March 12, 2016, for epidemiological reports on Zika virus outbreaks using the search terms \"Zika\" AND \"Basic reproductive number\". We applied no date or language restrictions. Our search identified one previous paper assessing the basic reproductive number, R0 of Zika virus in Yap Island, Federal State of Micronesia and in French Polynesia, but no papers estimating R0 using data from the Latin American Zika outbreak. Because of the sparsity of the data, we could not do a detailed systematic review at this point in time.\n\nAdded value of this studyWe report detailed epidemiological data on outbreaks in San Andres and Girardot, Colombia. Because such reports are currently unavailable, we provide early information on age and gender effects and the functioning of local and national surveillance in the second-most affected country in this epidemic. We provide early estimates of R0. Our results can be used by mathematical modelers to understand the future impact of the disease and potential spread.\n\nImplications of all the available evidenceWe report attack rates similar to those reported in the Yap Island outbreak. We find that Zika impacts individuals of all ages, though the most affected age group is 20 to 49 years of age. The surveillance system detected more cases among women in both areas, though this finding may be attributable to reporting bias. Our estimates of R0 imply that Zika has the capacity for widespread transmission in areas with the vector.

Epidemiology

The Mathematics of Descriptions in Descriptive Epidemiology

Identifying patterns of disease distribution in a population is the task of descriptive epidemiology, where the patterns are descriptions of how the disease is distributed in the population. In everyday public health practice it is descriptive epidemiology that is hard at work when we inform the public about health problems, monitor its health status, gather information for planning and administrative purposes, or generate hypotheses to be tested by further analytic investigations. In this paper we pursue a formalization of the population descriptions that are the heart of descriptive epidemiology. We call attention to a fact, so far unrecognized by epidemiologists, that the set of (complete) population descriptions, suitably defined, has an associated relational and algebraic structure. This formalization allows us to access mathematical properties of this structure of an epidemiologic dataset with methods developed over the past 30 years.

epidemiology

Whole genome analysis of local Kenyan and global sequences unravels the epidemiological and molecular evolutionary dynamics of RSV genotype ON1 strains

The respiratory syncytial virus (RSV) group A variant with the 72-nucleotide duplication in the G gene, genotype ON1, was first detected in Kilifi in 2012 and has almost completely replaced previously circulating genotype GA2 strains. This replacement suggests some fitness advantage of ON1 over the GA2 viruses, and might be accompanied by important genomic substitutions in ON1 viruses. Close observation of such a new virus introduction over time provides an opportunity to better understand the transmission and evolutionary dynamics of the pathogen. We have generated and analyzed 184 RSV-A whole genome sequences (WGS) from Kilifi (Kenya) collected between 2011 and 2016, the first ON1 genomes from Africa and the largest collection globally from a single location. Phylogenetic analysis indicates that RSV-A transmission into this coastal Kenya location is characterized by multiple introductions of viral lineages from diverse origins but with varied success in local transmission. We identify signature amino acid substitutions between ON1 and GA2 viruses within genes encoding the surface proteins (G, F), polymerase (L) and matrix M2-1 proteins, some of which were identified as positively selected, and thereby provide an enhanced picture of RSV-A diversity. Furthermore, five of the eleven RSV open reading frames (ORF) (i.e. G, F, L, N and P), analyzed separately, formed distinct phylogenetic clusters for the two genotypes. This might suggest that coding regions outside of the most frequently studied G ORF play a role in the adaptation of RSV to host populations with the alternative possibility that some of the substitutions are nothing more than genetic hitchhikers. Our analysis provides insight into the epidemiological processes that define RSV spread, highlights the genetic substitutions that characterize emerging strains, and demonstrates the utility of large-scale WGS in molecular epidemiological studies.\n\nAuthor summaryRespiratory syncytial virus (RSV) is the leading viral cause of severe pneumonia and bronchiolitis among infants and children globally. No vaccine exists to date. The high genetic variability of this RNA virus, characterized by group (A or B), genotype (within group) and variant (within genotype) replacement in populations, may pose a challenge to effective vaccine design by enabling immune response escape. To date most sequence data exists for the highly variable G gene encoding the RSV attachment protein, and there is little globally-sampled RSV genomic data to provide a fine resolution of the epidemiology and evolutionary dynamics of the pathogen. Here we use long-term RSV surveillance in coastal Kenya to track the introduction, spread and evolution of a new RSV genotype known as ON1 (having a 72-nucleotide duplication in the G gene). We present a set of 184 RSV-A whole genomes, including 176 of RSV ON1 (the first from Africa), describe patterns of local ON1 spread and show genome-wide changes between the two major RSV-A genotypes that may define the pathogens adaptation to the host. These findings have implications for vaccine design and improved understanding of RSV epidemiology and evolution.

epidemiology

Assessing the accuracy of Approximate Bayesian Computation approaches to infer epidemiological parameters from phylogenies

Phylodynamics typically rely on likelihood-based methods to infer epidemiological parameters from dated phylogenies. These methods are essentially based on simple epidemiological models because of the difficulty in expressing the likelihood function analytically. Computing this function numerically raises additional challenges, especially for large phylogenies. Here, we use Approximate Bayesian Computation (ABC) to circumvent these problems. ABC is a likelihood-free method of parameter inference, based on simulation and comparison between target data and simulated data, using summary statistics. We simulated target trees under several epidemiological scenarios in order to assess the accuracy of ABC methods for inferring epidemiological parameter such as the basic reproduction number (R0), the mean duration of infection, and the effective host population size. We designed many summary statistics to capture the information in a phylogeny and its corresponding lineage-through-time plot. We then used the simplest ABC method, called rejection, and its modern derivative complemented with adjustment of the posterior distribution by regression. The availability of machine learning techniques including variable selection, motivated us to compute many summary statistics on the phylogeny. We found that ABC-based inference reaches an accuracy comparable to that of likelihood-based methods for birth-death models and can even outperform existing methods for more refined models and large trees. By re-analysing data from the early stages of the recent Ebola epidemic in Sierra Leone, we also found that ABC provides more realistic estimates than the likelihood-based methods, for some parameters. This work shows that the combination of ABC-based inference using many summary statistics and sophisticated machine learning methods able to perform variable selection is a promising approach to analyse large phylogenies and non-trivial models.

Bioinformatics

Use of whole-genome sequencing in the epidemiology of Campylobacter jejuni infections: state-of-knowledge

High-throughput whole-genome sequencing (WGS) is a revolutionary tool in public health microbiology and is gradually substituting classical typing methods in surveillance of infectious diseases. In combination with epidemiological methods, WGS is able to identify both sources and transmission-pathways during disease outbreak investigations. This review provides the current state of knowledge on the application of WGS in the epidemiology of Campylobacter jejuni, the leading cause of bacterial gastroenteritis in the European Union. We describe how WGS has improved surveillance and outbreak detection of C. jejuni infections and how WGS has increased our understanding of the evolutionary and epidemiological dynamics of this pathogen. However, the full implementation of this methodology in real-time is still hampered by a few hurdles. The limited insight into the genetic diversity of different lineages of C. jejuni impedes the validity of assumed genetic relationships. Furthermore, efforts are needed to reach a consensus on which analytic pipeline to use and how to define the strains cut-off value for epidemiological association while taking the needs and realities of public health microbiology in consideration. Even so, we claim that ample evidence is available to support the benefit of integrating WGS in the monitoring of C. jejuni infections and outbreak investigations.

Microbiology

Evolution of subgenomic RNA shapes dengue virus adaptation and epidemiological fitness

Genetic changes in the dengue virus (DENV) genome affects viral fitness both clinically and epidemiologically. Even in the 3 untranslated region (3UTR), mutations could impact the formation of subgenomic flaviviral RNA (sfRNA) and the specificity of sfRNA in inhibiting host proteins necessary for successful viral replication. Indeed, we have recently shown that mutations in the 3UTR of DENV2 affected its ability to inhibit TRIM25 E3 ligase activity to reduce interferon (IFN) expression, which potentially contributed to the emergence of a new viral clade during the 1994 dengue epidemic in Puerto Rico. However, whether differences in 3UTRs shaped DENV evolution on a larger scale remains incompletely understood. Herein, we combined RNA phylogeny with phylogenetics to gain insights on sfRNA evolution. We found that sfRNA structures are under purifying selection and highly conserved despite sequence divergence. Interestingly, only the second flaviviral Nuclease-resistant RNA (fNR2) structure of DENV-2 has undergone strong positive selection. Epidemiological reports also suggest that nucleotide substitutions in fNR2 may drive DENV-2 epidemiological fitness, possibly through sfRNA-protein interactions. Collectively, our findings indicate that 3UTRs are important determinants of DENV fitness in human-mosquito cycles.\n\nHighlightsO_LIDengue viruses (DENV) preserve RNA elements in their 3 untranslated region (UTR).\nC_LIO_LISite-specific quantification of natural selection revealed positive selection on DENV2 sfRNA.\nC_LIO_LIFlaviviral nuclease-resistant RNA (fNR) structures in DENV 3UTRs contribute to DENV speciation.\nC_LIO_LIA highly evolving fNR structure appears to increase DENV-2 epidemiological fitness.\nC_LI

microbiology

Core Genome Multi Locus Sequence Typing and Single Nucleotide Polymorphism Analysis in the Epidemiology of Brucella melitensis Infections

The use of whole genome sequencing (WGS) using next generation sequencing (NGS) technology has become a widely accepted method for microbiology laboratories in the application of molecular typing for outbreak tracing and genomic epidemiology. Several studies demonstrated the usefulness of WGS data analysis through Single Nucleotide Polymorphism (SNP) calling from a reference sequence analysis for Brucella melitensis, whereas gene-by-gene comparison through core-genome Multilocus Sequence Typing (cgMLST) has not been explored so far. The current study developed an allele-based method cgMLST and compared its performance to the genome-wide SNP approach and the traditional MLVA on a defined sample collection. The dataset comprised of 37 epidemiologically linked animal cases of brucellosis as well as 71 epidemiologically unrelated human and animal isolates collected in Italy. The cgMLST scheme generated in this study contained 2,687 targets of the B. melitensis 16M reference genome (75.4% of the complete genome). We established the potential criteria necessary for inclusion of an isolate into a brucellosis outbreak cluster to be [&le;]4 loci in the cgMLST and [&le;]10 in WGS SNP analysis. CgMLST and SNP analysis provided much higher phylogenetic distance resolution than MLVA, particularly for strains belonging to the same lineage thus allowing diverse and unrelated genotypes to be identified with greater confidence. The application of this cgMLST scheme to the characterization of B. melitensis strains provided insights into the epidemiology of this pathogen and it is a candidate to be a benchmark tool for outbreak investigations in human and animal brucellosis.

microbiology

Genotyping and epidemiological metadata provides new insights into population structure of Xanthomonas isolated from walnut trees

Xanthomonas arboricola pv. juglandis (Xaj) is the etiological agent of walnut diseases affecting leaves, fruits, branches and trunks. Although this phytopathogen is widely spread in walnut producing regions and has a considerable genetic diversity, there is still a poor understanding of its epidemic behaviour. To shed some light on the epidemiology of these bacteria, 131 Xanthomonas isolates obtained from 64 walnut trees were included in this study considering epidemiological metadata such as year of isolation, bioclimatic regions, walnut cultivars, production regimes, host walnut specimen and plant organs. Genetic diversity was assessed by multilocus sequence analysis (MLSA) and dot blot hybridization patterns obtained with nine Xaj-specific DNA markers (XAJ1 - XAJ9). The results showed that Xanthomonas isolates grouped in ten distinct MLSA clusters and in 18 hybridization patterns (HP). The majority of isolates (112 out of 131) were closely related with X. arboricola strains of pathovar juglandis as revealed by MLSA (clusters I to VI) and hybridize with more than five Xaj-specific markers. Nineteen isolates clustered in four MLSA groups (clusters VII to X) which do not include Xaj strains, and hybridize to less than five markers. Taking this data together, was possible to distinguish 17 lineages of Xaj, three lineages of X. arboricola and 11 lineages of Xanthomonas sp. Some Xaj lineages appeared to be widely distributed and prevalent across the different bioclimatic regions and apparently not constrained by the other features considered. Assessment of type III effector genes and pathogenicity tests revealed that representative lineages of MLSA clusters VII to X were nonpathogenic on walnut, with exception for strain CPBF 424, making this bacterium particularly appealing to address Xanthomonas pathoadaptations to walnut.\n\nIMPORTANCEXanthomonas arboricola pv. juglandis is one of the most serious threats of walnut trees. New disease epidemics caused by this phytopathogen has been a big concern causing high economic losses on walnut production worldwide. Using a comprehensive sampling methodology to disclose the diversity of walnut infective Xanthomonas, we were able to identify a genetic diversity higher than previously reported and generally independent of bioclimatic regions and the other epidemiological features studied. Furthermore, co-colonization of the same plant sample by distinct Xanthomonas strains were frequent and suggested a sympatric lifestyle. The extensive sampling carried out resulted in a set of non-arboricola Xanthomonas sp. strains, including a pathogenic strain, therefore diverging from the nonpathogenic phenotype that have been associated to these atypical strains, generally considered to be commensal. This new strain might be particularly informative to elucidate novel pathogenicity traits and unveil pathogenesis evolution within walnut infective xanthomonads. Beyond extending the present knowledge about walnut infective xanthomonads, this study might contribute to provide a methodological framework for phytopathogen epidemiological studies, still largely disregarded.

microbiology

Sample size calculation for estimating key epidemiological parameters using serological data and mathematical modelling

BackgroundOur work was motivated by the need to, given serum availability and/or financial resources, decide on which samples to test for different pathogens in a serum bank. Simulation-based sample size calculations were performed to determine the age-based sampling structures and optimal allocation of a given number of samples for testing across various age groups best suited to estimate key epidemiological parameters (e.g., seroprevalence or force of infection) with acceptable precision levels in a cross-sectional seroprevalence survey.\n\nMethodsStatistical and mathematical models and three age-based sampling structures (survey-based structure, population-based structure, uniform structure) were used. Our calculations are based on Belgian serological survey data collected in 2002 where testing was done, amongst others, for the presence of IgG antibodies against measles, mumps, and rubella, for which a national mass immunisation programme was introduced in 1985 in Belgium, and against varicella-zoster virus and parvovirus B19 for which the endemic equilibrium assumption is tenable in Belgium.\n\nResultsThe optimal age-based sampling structure to use in the sampling of a serological survey as well as the optimal allocation distribution varied depending on the epidemiological parameter of interest for a given infection and between infections.\n\nConclusionsWhen estimating key epidemiological parameters with acceptable levels of precision within the context of a single cross-sectional serological survey, attention should be given to the age-based sampling structure. Simulation-based sample size calculations in combination with mathematical modelling can be utilised for choosing the optimal allocation of a given number of samples over various age groups.

epidemiology

Estimating the current burden of Chagas disease in Mexico: a systematic review of epidemiological surveys from 2006 to 2017

BackgroundIn Mexico, estimates of Chagas disease prevalence and burden vary widely. Updating surveillance data is therefore an important priority to ensure that Chagas disease does not remain a barrier to the development of Mexicos most vulnerable populations.\n\nMethodology/Principal FindingsThe aim of this systematic review was to analyze the literature on epidemiological surveys to estimate Chagas disease prevalence and burden in Mexico, during the period 2006 to 2017. A total of 2,764 articles were screened and 38 were retained for the final analysis. Epidemiological surveys have been performed in most of Mexico, but with variable geographic coverage. Based on studies reporting confirmed cases, i.e. using at least 2 serological tests, the national seroprevalence of Trypanosoma cruzi infection was 2.26% [95% Confidence Interval (CI) 2.12-2.41], suggesting that there are 2.71 million cases in Mexico. Studies focused on pregnant women, which may transmit the parasite to their newborn during pregnancy, reported a seroprevalence of 1.00% [95% CI 0.87-1.14], suggesting that there are 22,930 births from T. cruzi infected pregnant women per year, and 1,445 cases of congenitally infected newborns per year. Children under 18 years had a seropositivity rate of 1.49% [95% CI 1.20-1.85], which indicate ongoing transmission. Finally, cases of T. cruzi infection in blood donors have also been reported in most states, with a national seroprevalence of 0.51% [95% CI 0.49-0.53].\n\nConclusions/SignificanceOur analysis suggests a disease burden for T. cruzi infection higher than previously recognized, highlighting the urgency of establishing Chagas disease surveillance and control as a key national public health priority in Mexico, to ensure that it does not remain a major barrier to the economic and social development of the countrys most vulnerable populations.\n\nAuthor summaryIn Mexico, estimates of Chagas disease prevalence and burden vary widely due to the ecology and epidemiology of this disease resulting of many geographical, ecological, biological, and social interactions. Better data are thus urgently needed to help develop appropriate public health programs for disease control and patient care. In this study we analyzed published data on T. cruzi seroprevalence infection in Mexico between 2006 and 2017. This systematic review shows a national seroprevalence of T. cruzi infection of 2.26% [95%CI 2.12-2.41], with over 2.71 million cases in Mexico, which is higher than previously recognized. The presence of T. cruzi infection in specific subpopulations such as pregnant women, children and blood donors also informs on specific risks of infection and call for the implementation of well-established control interventions. This work confirms the place of Mexico as the country with the largest number of cases, highlighting the urgency of establishing Chagas disease control as a key national public health priority.

epidemiology

Genomic epidemiology and global diversity of the emerging bacterial pathogen Elizabethkingia anophelis

Elizabethkingia anophelis is an emerging pathogen. Genomic analysis of strains from clinical, environmental or mosquito sources is needed to understand the epidemiological emergence of E. anophelis and to uncover genetic elements implicated in antimicrobial resistance, pathogenesis, or niche adaptation. Here, the genomic sequences of two nosocomial isolates that caused neonatal meningitis in Bangui, Central African Republic, were determined and compared with Elizabethkingia isolates from other world regions and sources. Average nucleotide identity firmly confirmed that E. anophelis, E. meningoseptica and E. miricola represent distinct genomic species and led to re-identification of several strains. Phylogenetic analysis of E. anophelis strains revealed several sublineages and demonstrated a single evolutionary origin of African clinical isolates, which carry unique antimicrobial resistance genes acquired by horizontal transfer. The Elizabethkingia genus and the species E. anophelis had pan-genomes comprising respectively 7,801 and 6,880 gene families, underlining their genomic heterogeneity. African isolates were capsulated and carried a distinctive capsular polysaccharide synthesis cluster. A core-genome multilocus sequence typing scheme applicable to all Elizabethkingia isolates was developed, made publicly available (http://bigsdb.web.pasteur.fr/elizabethkingia), and shown to provide useful insights into E. anophelis epidemiology. Furthermore, a clustered regularly interspaced short palindromic repeats (CRISPR) locus was uncovered in E. meningoseptica, E. miricola and in a few E. anophelis strains. CRISPR spacer variation was observed between the African isolates, illustrating the value of CRISPR for strain subtyping. This work demonstrates the dynamic evolution of E. anophelis genomes and provides innovative tools for Elizabethkingia identification, population biology and epidemiology.\n\nIMPORTANCEElizabethkingia anophelis is a recently recognized bacterial species involved in human infections and outbreaks in distinct world regions. Using whole-genome sequencing, we showed that the species comprises several sublineages, which differ markedly in their genomic features associated with antibiotic resistance and host-pathogen interactions. Further, we have devised high-resolution strain subtyping strategies and provide an open genomic sequence analysis tool, facilitating the investigation of outbreaks and tracking of strains across time and space. We illustrate the power of these tools by showing that two African healthcare-associated meningitis cases observed 5 years apart were caused by the same strain, providing evidence that E. anophelis can persist in the hospital environment.

Microbiology

Microevolution of Bank Voles (Myodes glareolus) at Neutral and Immune Related Genes During Multiannual Dynamic Cycles: Consequences for Puumala hantavirus Epidemiology

Understanding how host dynamics, including variations of population size and dispersal, may affect the epidemiology of infectious diseases through ecological and evolutionary processes is an active research area. Here we focus on a bank vole (Myodes glareolus) metapopulation surveyed in Finland between 2005 and 2009. Bank vole is the reservoir of Puumala hantavirus (PUUV), the agent of nephropathia epidemica (NE, a mild form of hemorrhagic fever with renal symptom) in humans. M glareolus populations experience multiannual density fluctuations that may influence the level of genetic diversity maintained in bank voles, PUUV prevalence and NE occurrence. We examine bank vole metapopulation genetics at presumably neutral markers and immune-related genes involved in susceptibility to PUUV (Tnf-promoter, Mhc-Drb, Tlr4, Tlr7 and Mx2 gene) to investigate the links between population dynamics, microevolutionary processes and PUUV epidemiology. We show that genetic drift slightly and transiently affects neutral and adaptive genetic variability within the metapopulation. Gene flow seems to counterbalance its effects during the multiannual density fluctuations. The low abundance phase may therefore be too short to impact genetic variation in the host, and consequently viral genetic diversity. Environmental heterogeneity does not seem to affect vole gene flow, which might explain the absence of spatial structure previously detected in PUUV in this area. Besides, our results suggest the role of vole dispersal on PUUV circulation through sex-specific and density-dependent movements. We find little evidence of selection acting on immune-related genes within this metapopulation. Footprint of positive selection is detected at Tlr-4 gene in 2008 only. We observe marginally significant associations between Mhc-Drb haplotypes and PUUV serology, and between Mx2 genotype and PUUV genogroups. These results show that microevolutionary changes and PUUV epidemiology in this metapopulation are mainly driven by neutral processes, although the relative effects of neutral and adaptive forces could vary temporally with density fluctuations.

Evolutionary Biology

Beta regression improves the detection of differential DNA methylation for epigenetic epidemiology

BackgroundDNA methylation is the most readily assayed epigenetic mark, possessing confirmed relationships with gene expression, imprinting, and chromatin accessibility.Given the increasingly widespread use of DNA methylation microarrays in population-scale epidemiological applications, we sought to determine which methods provided the greatest statistical power to reproducibly detect differences in DNA methylation across various conditions,using publicly available data sets on tissue type and aging.\n\nResultsBeta regression, as proposed originally by Ferrari and Cribari-Neto, yielded more validated hits in each of our comparisons than any other method under consideration, both in a regression setting and in comparisons to two-group tests such as the Wilcoxon-Mann-Whitney, Student t, and Welch t tests.In large cohorts of whole blood samples, we corrected for compositional differences and batch effects, and found that marginal likelihood ratio tests from beta regression models uniformly dominate popular alternatives based on linear models.The superior sensitivity and specificity exhibited by beta regression in epidemiologically relevant cohort sizes corresponded to approximately a 2% increase in sensitivity at the same specificity when compared to linear models fitted on raw beta values (proportion of signal intensity due to the methylated allele), M-values, or rankquantile normalized values.\n\nConclusionsInvestigators should consider beta regression to maximize statistical power in studies of DNA methylation using microarrays.At epidemiologically relevant sample sizes, with typical quality control procedures (compositional and batch effect correction), cross-cohort agreement uniformly favors beta regression over popular alternatives.

Bioinformatics

Integrating Patient And Whole Genome Sequencing Data To Provide Insights Into The Epidemiology Of Seasonal Influenza A(H3N2) Viruses

Genetic surveillance of seasonal influenza is largely focused upon sequencing of the haemagglutinin gene. Consequently, our understanding of the contribution of the remaining seven gene segments to the evolution and epidemiological dynamics of seasonal influenza is relatively limited. The increased availability of next generation sequencing technologies allows rapid and economic whole genome sequencing (WGS). Here, 150 influenza A(H3N2) positive clinical specimens with linked epidemiological data, from the 2014/15 season in Scotland, were sequenced directly using both Sanger sequencing of the HA1 region and WGS using the Illumina MiSeq platform. Sequences generated by both methods were highly consistent and WGS provided on average >90% whole genome coverage. As reported in other European countries during 2014/15, all strains belonged to genetic group 3C, with subgroup 3C.2a predominating. Inter-subgroup reassortants were identified (9%), including three 3C.3 viruses descended from a single reassortment event, which had persisted in the population. Significant phylogenetic associations with cases of severe acute respiratory illness observed herein warrant further investigation. Severe cases were also more likely to be associated with reassortant viruses (odds ratio: 4.4 (1.3-15.5)) and occur later in the season. These results suggest that increased levels of WGS, linked to clinical and epidemiological data, could improve influenza surveillance.

microbiology