Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Epidemic Potential for Human Infection with Influenza A (H7N9) Virus in China through Web Search Behaviors: A Data-Driven Study

Since the beginning of September 2016, a steep upsurge of the human cases of avian influenza A (H7N9) virus has been reported in China, which are alarming public concern for the pandemic potential of the H7N9 virus. In this study, we collected the data from H7N9 epidemics and H7N9-related Baidu Search Index (BSI) in China between September 2013 and June 2017. And we observed a strong correlation between the numbers of Influenza A (H7N9) cases and H7N9-related BSI in Guangdong province and Shanghai municipality (p<0.001). Autoregressive integrated moving average (ARIMA) models were constructed for the dynamic estimation of seasonal H7N9 outbreaks in 2016-2017 and the online search data acted as an external regressor with the historical H7N9 epidemic data in the forecasting model to improve the quality of predictions. Predictions by the models closely matched the actual numbers of reported cases during current H7N9 epidemic season. Especially, the estimated numbers of reported cases sharply increased to reach 49.88 (95% CI: 0-194.05) in Guangdong and 9.05 (95% CI: 0-37.43) in Shanghai from December 2016 to June 2017. Moreover, this accessible and flexible dynamic forecast model could be used in the monitoring of H7N9 virus to provide advanced warning of future emerging infection diseases.\n\nAuthor summaryAs the availability and popularity of the internet has greatly increased in recent years, an increasing number of cyber users, including patients and their family members, search online for health information on personal computers (PCs) and mobile phones (MPs) before seeking medical attention, making it possible to investigate the influenza prevalence by monitoring changes in frequencies of uses of particular search terms. In this study, we collected the data from H7N9 epidemics and H7N9-related Baidu Search Index (BSI) in China between September 2013 and June 2017. And then, we showed a strong correlation between the numbers of Influenza A (H7N9) cases and H7N9-related BSI in Guangdong province and Shanghai municipality (p<0.001). Furthermore, we reconstructed an improved dynamic forecasting method for outbreaks of H7N9 influenza using Autoregressive integrated moving average (ARIMA) models to predict future patterns of H7N9 transmission and the online search data acted as an external regressor with the historical H7N9 epidemic data in the forecasting model to improve the quality of predictions. Our results suggest that data from the Baidu search engine, combed with data from a traditional disease surveillance system, may be considered for early detection of H7N9 influenza outbreaks in mainland China.

epidemiology

Conjunction of Factors Triggering Waves of Seasonal Influenza

Understanding the subtle confluence of factors triggering pan-continental, seasonal epidemics of influenza-like illness is an extremely important problem, with the potential to save tens of thousands of lives and billions of dollars every year in the US alone. Beginning with several large, longitudinal datasets on putative factors and clinical data on the disease and health status of over 150 million human subjects observed over a decade, we investigated the source and the mechanistic triggers of epidemics. Our analysis included insurance claims for a significant cross-section of the US population in the past decade, human movement patterns inferred from billions of tweets, whole-US weekly weather data covering the same time span as the medical records, data on vaccination coverage over the same period, and sequence variations of key viral proteins. We also explicitly accounted for the spatio-temporal auto-correlations of infectious waves, and a host of socioeconomic and demographic factors. We carried out multiple orthogonal statistical analyses on these diverse, large geo-temporal datasets to bolster and corroborate our findings. We conclude that the initiation of a pan-continental influenza wave emerges from the simultaneous realization of a complex set of conditions, the strongest predictor groups are as follows, ranked by importance: (1) the host populations socio- and ethno-demographic properties; (2) weather variables pertaining to relevant area specific humidity, temperature, and solar radiation; (3) the virus antigenic drift over time; (4) the host populations land-based travel habits, and; (5) the spatio-temporal dynamics immediate history, as reflected in the influenza wave autocorrelation. The models we infer are demonstrably predictive (area under the Receiver Operating Characteristic curve {approx} 80%) when tested with out-of-sample data, opening the door to the potential formulation of new population-level intervention and mitigation policies.

epidemiology

Identifying the dynamic contact network of infectious disease spread

The spread of pathogens fundamentally depends on the underlying contacts between individuals. Modeling infectious disease dynamics through contact networks is sometimes challenging, however, due to a limited understanding of pathogen transmission routes and infectivity. We developed a novel tool, INoDS (Identifying Network models of infectious Disease Spread) that estimates the predictive power of empirical contact networks to explain observed patterns of infectious disease spread. We show that our method is robust to partially sampled contact networks, incomplete disease information, and enables hypothesis testing on transmission mechanisms. We demonstrate the applicability of our method in two host-pathogen systems: Crithidia bombi in bumble bee colonies and Salmonella in wild Australian sleepy lizard populations. The performance of INoDS in synthetic and complex empirical systems highlights its role in identifying transmission pathways of novel or neglected pathogens, as an alternative approach to laboratory transmission experiments, and overcoming common data-collection constraints.

epidemiology

Social fluidity mobilizes infectious disease in human and animal populations

Humans and other group-living animals tend to distribute their social effort disproportionately. Individuals predominantly interact with a small number of close companions while maintaining weaker social bonds with less familiar group members. By incorporating this behaviour into a mathematical model we find that a single parameter, which we refer to as social fluidity, controls the rate of social mixing within the group. We compare the social fluidity of 13 species by applying the model to empirical human and animal social interaction data. To investigate how social behavior influences the likelihood of an epidemic outbreak we derive an analytical expression of the relationship between social fluidity and the basic reproductive number of an infectious disease. For highly fluid social behaviour disease transmission is revealed to be density-dependent. For species that form more stable social bonds, the model describes frequency-dependent transmission that is sensitive to changes in social fluidity.

epidemiology

Centenarian Hotspots in Denmark

BackgroundThe study of regions with high prevalence of centenarians is motivated by a desire to find determinants of healthy ageing. While existing research has focused on selected candidate regions, we explore the existence of hotspots in the whole the Denmark, which is a small and homogeneous country.\n\nMethodsWe performed a Kulldorff spatial scan across the whole of Denmark, searching for regions of birth and regions of residence at age 71 where a significantly increased percentage of the cohort born 1906-1915 became centenarians. We compared mortality hazards for the hotspot regions to the rest of the country by sex and residence at age 71.\n\nResultsWe found a birth hotspot of 222 centenarians, 1.37 times more than the expected number, centered on a group of fairly remote rural islands. Mortality was lower for those born in the hotspot, and the advantage was strongest in those born and remaining in the hotspot. At age 71, we found one primary hotspot with 1.39 times the expected number of centenarians, located in the generally high-income suburbs of the Danish capital.\n\nConclusionWe identified a Danish centenarian birth hotspot that persisted over a period of at least 30 years. Exactly what drives the appearance of this birth hotspot is unknown, but selection may explain hotspots based on residence at age 71.

epidemiology

Climate change drives uncertain global shifts in potential distribution and seasonal risk of Aedes-transmitted viruses

AbstractForecasting the impacts of climate change on Aedes-borne viruses--especially dengue, chikungunya, and Zika--is a key component of public health preparedness. We apply an empirically parameterized Bayesian transmission model of Aedes-borne viruses for the two vectors Aedes aegypti and Ae. albopictus as a function of temperature to predict cumulative monthly global transmission risk in current climates, and compare with projected risk in 2050 and 2080 based on general circulation models (GCMs). Our results show that if mosquito range shifts track optimal temperatures for transmission (26-29 {degrees}C), we can expect poleward shifts in Aedes-borne virus distributions. However, the differing thermal niches of the two vectors produce different patterns of shifts under climate change. More severe climate change scenarios produce proportionally worse population exposures from Ae. aegypti, but not from Ae. albopictus in the most extreme cases. Expanding risk of transmission from both mosquitoes will likely be a serious problem, even in the short term, for most of Europe; but significant reductions are also expected for Aedes albopictus, most noticeably in southeast Asia and west Africa. Within the next century, nearly a billion people are threatened with new exposure to both Aedes spp. in the worst-case scenario; but massive net losses in risk are noticeable for Ae. albopictus, especially in terms of year-round transmission, marking a global shift towards more seasonal risk across regions. Many other complicating factors (like mosquito range limits and viral evolution) exist, but overall our results indicate that while climate change will lead to both increased and new exposures to vector-borne disease, the most extreme increases in Ae. albopictus transmission are predicted to occur at intermediate climate change scenarios.\n\nAuthor SummaryThe established scientific consensus indicates that climate change will severely exacerbate the risk and burden of Aedes-transmitted viruses, including dengue, chikungunya, Zika, West Nile virus, and other significant threats to global health security. Here, we show that the story is more complicated, first and foremost due to differences between the more heat-tolerant Aedes aegypti and the more heat-limited Ae. albopictus. Almost a billion people could face their first exposure to viral transmission from either mosquito in the worst-case scenario, especially in Europe and high-elevation tropical and subtropical regions. On the other hand, while year-round transmission potential from Ae. aegypti is likely to expand (especially in south Asia and sub-Saharan Africa), Ae. albopictus loses significant ground in the tropics, marking a global shift towards seasonal risk as the tropics eventually become too hot for transmission by Ae. albopictus. Complete mitigation of climate change to a pre-industrial baseline could protect almost a billion people from arbovirus range expansions; but middle-of-the-road mitigation may actually produce the greatest expansion in the potential for viral transmission by Ae. albopictus. In any scenario, mitigating climate change also shifts the burden of both dengue and chikungunya (and potentially other Aedes transmitted viruses) from higher-income regions back onto the tropics, where transmission might otherwise start to be curbed by rising temperatures.

epidemiology

The Impact of Education on Myopia: A bidirectional Mendelian randomisation analysis in UK Biobank

Myopia, or short-sightedness, is one of the leading causes of visual disability in the World. The prevalence of myopia has risen steadily over recent decades, reaching epidemic levels in Southeast Asia. Observational studies have reported associations between educational attainment and myopia. Whether education causes myopia, myopic children are more intelligent, or another factor, like higher socioeconomic status, causes both is unclear since observational studies are prone to confounding and randomised trials of education are unethical. Using bidirectional Mendelian Randomisation, a form of instrumental variable (IV) analysis free from confounding, we show that every additional year in education leads to an increase in myopic refractive error, but that myopia does not lead to higher educational attainment. Our results suggest that current educational methods contribute to the global burden of myopia, and argue that educational policies and practices should take account of this to reduce future visual disability in the population.

epidemiology

Cadmium Exposure Increases The Risk Of Juvenile Obesity: A Human And Zebrafish Comparative Study

OBJECTIVEHuman obesity is a complex metabolic disorder disproportionately affecting people of lower socioeconomic strata, and ethnic minorities, especially African Americans and Hispanics. Although genetic predisposition and a positive energy balance are implicated in obesity, these factors alone do not account for the excess prevalence of obesity in lower socioeconomic populations. Therefore, environmental factors, including exposure to pesticides, heavy metals, and other contaminants, are agents widely suspected to have obesogenic activity, and they also are spatially correlated with lower socioeconomic status. Our study investigates the causal relationship between exposure to the heavy metal, cadmium (Cd), and obesity in a cohort of children and a zebrafish model of adipogenesis.\n\nDESIGNAn extensive collection of first trimester maternal blood samples obtained as part of the Newborn Epigenetics Study (NEST) were analyzed for the presence Cd, and these results were cross analyzed with the weight-gain trajectory of the children through age five years. Next, the role of Cd as a potential obesogen was analyzed in an in vivo zebrafish model.\n\nRESULTSOur analysis indicates that the presence of Cd in maternal blood during pregnancy is associated with increased risk of juvenile obesity in the offspring, independent of other variables, including lead (Pb) and smoking status. Our results are recapitulated in a zebrafish model, in which exposure to Cd at levels approximating those observed in the NEST study is associated with increased adiposity.\n\nCONCLUSIONOur findings identify Cd as potential human obesogen. Moreover, these observations are recapitulated in a zebrafish model, suggesting that the underlying mechanisms may be evolutionarily conserved, and that zebrafish may be a valuable model for uncovering pathways leading to Cd-mediated obesity in human populations.

epidemiology

Stacked Generalization: An Introduction to Super Learning

Stacked generalization is an ensemble method that allows researchers to combine several different prediction algorithms into one. Since its introduction in the early 1990s, the method has evolved several times into what is now known as \"Super Learner\". Super Learner uses V -fold cross-validation to build the optimal weighted combination of predictions from a library of candidate algorithms. Optimality is defined by a user-specified objective function, such as minimizing mean squared error or maximizing the area under the receiver operating characteristic curve. Although relatively simple in nature, use of the Super Learner by epidemiologists has been hampered by limitations in understanding conceptual and technical details. We work step-by-step through two examples to illustrate concepts and address common concerns.

epidemiology

A New Mechanism for Cost Savings in NHS Prescribing: Minimising "Price-Per-Unit"

BackgroundMinimising prescription costs while maintaining quality is a core element of delivering high value healthcare. There are various strategies to achieve savings, but almost no research to date on determining the most effective approach. We describe a new method of identifying potential savings due to large national variations in drug cost, including variation in generic drug cost; and compare these with potential savings from an established method (generic prescribing).\n\nMethodsWe used English NHS Digital prescribing data, from October 2015 to September 2016. Potential cost savings were calculated by determining the price-per-unit (e.g. pill, ml) for each drug and dose within each general practice. This was compared against the same cost for the practice at the lowest cost decile, to determine achievable savings. We compared these price-per-unit savings to the savings possible from generic switching; and determined the chemicals with the highest savings nationally. A senior pharmacist manually assessed whether a random sample of savings were practically achievable.\n\nResultsWe identified a theoretical maximum of {pound}410M of savings over 12 months. {pound}273M of these savings were for individual prescribing changes worth over {pound}50 per practice per month; this compares favorably with generic switching, where only {pound}35M of achievable savings were identified. The biggest savings nationally were on glucose blood testing reagents ({pound}12M), fluticasone propionate ({pound}9M) and venlafaxine ({pound}8M). Approximately half of all savings were deemed practically achievable.\n\nDiscussionWe have developed a new method to identify and enable large potential cost savings within NHS community prescribing. Given the current pressures on the NHS, it is vital that these potential savings are realised. Our tool enabling doctors to achieve these savings is now launched in pilot form. However savings could potentially be achieved more simply through national policy change.\n\nAbbreviations

epidemiology

Identifying factors that may improve mechanistic forecasting models for influenza

Influenza causes substantial morbidity and mortality and places strain on healthcare systems, some of which could be mitigated by accurate forecasting. Specific humidity and school vacations have both been shown independently to affect the transmission dynamics of influenza at large spatial scales. Here, we compare the ability of five compartmental transmission models, which include these two processes, to explain influenza-like-illness (ILI) incidence data for five United States counties for which school vacations and specific humidity data were available over a span of four seasons. We used the models in two different ways. First we fitted all available data at the same time and assessed model performance using standard measures of parsimony and goodness-of-fit. Then we conducted a retrospective forecasting study in which we attempted to predict incidence beyond a given week by fitting to data available up to that week. In general, when fitting the data using the whole season, we found that either specific humidity, school closures, or a combined model incorporating both effects captured the variability in incidence better than a fully constrained SIR-like model. Moreover, where these factors play a role, the timing of the variations suggests a causal relationship. When school vacations and specific humidity were important, the model-estimated parameters were broadly consistent. Retrospective forecasting simulations were consistent with the explanatory use of the models, with both specific humidity and school vacations giving more accurate forecasts than a simple SIR-like model in some populations and for some seasons. Our results suggest that influenza forecast models should test for the importance of different factors such as school vacations and specific humidity on a population-by-population and year-by-year basis.\n\nAuthor summaryUnderstanding the underlying factors that contribute to the transmission of influenza is crucial for developing models with predictive capabilities. In this study, we address two key effects: humidity and school vacations. We show that both can play an important role, depending on the location of the population as well as the timing of the school vacations. We then demonstrate how such mechanistic models can be used to forecast an influenza season as it unfolds, including estimates of the uncertainty of the predictions at each week of the forecast. To make better influenza forecasts, models need to test whether factors such as school vacations and specific humidity are important for a given season and population.

epidemiology

On the relative role of different age groups during epidemics associated with the respiratory syncytial virus

BackgroundWhile RSV circulation results in high burden of hospitalization, particularly among infants, young children and the elderly, little is known about the role of different age groups in propagating annual RSV epidemics in the community.\n\nMethodsDuring a communicable disease outbreak, some subpopulations may play a disproportionate role during the outbreak's ascent due to increased susceptibility and/or contact rates. Such subpopulations can be identified by considering the proportion that cases in a subpopulation represent among all cases in the population occurring before (Bp) and after the epidemic peak (Ap) to calculate the subpopulation's relative risk, RR=Bp/Ap. We estimated RR for several age groups using data on RSV hospitalizations in the US between 2001-2012 from the Healthcare Cost and Utilization Project (HCUP).\n\nResultsChildren aged 3-4y and 5-6y each had the highest RR estimate for 5/11 seasons in the data, with RSV hospitalization rates in infants being generally higher during seasons when children aged 5-6y had the highest RR estimates. Children aged 2y had the highest RR estimate during one season. RR estimates in infants and individuals aged 11y and older were mostly lower than in children aged 1-10y.\n\nConclusionsThe RR estimates suggest that preschool and young school-age children have the leading relative roles during RSV epidemics. We hope that those results will aid in the design of RSV vaccination policies.

epidemiology

Predictors of intestinal inflammation in asymptomatic first-degree relatives of patients with Crohn’s disease

ObjectiveRelatives of individuals with Crohns disease (CD) carry an increased number of CD-associated genetic variants and are at increased risk of developing the disease. Multiple environmental and genetic factors contribute to this increased risk. We aimed to estimate the utility of genotype, smoking, family history, and a panel of biomarkers to predict risk in asymptomatic first-degree relatives (FDRs) of CD patients.\n\nDesignWe calculated a combined genotype (72 CD-associated genetic markers) and smoking relative risk score in 454 FDRs, and performed capsule endoscopy and collected 22 biomarkers in individuals from the highest and lowest risk quartiles. We then predicted small intestinal inflammation using genetic risk score, smoking status, number of relatives with CD, capsule transit time, and the panel of biomarkers in 124 individuals with complete data. Our principal analysis was to calculate the predictive utility from two machine learning classifiers: an elastic net and a random forest.\n\nResultsBoth classifiers successfully predicted FDRs with intestinal inflammation: elastic net (AUC=0.80, 95% CI: 0.62-0.98), random forest (AUC=0.87, 95% CI: 0.75-1.00). The elastic net selected a 3-predictor solution: CD family history (OR=1.31), genetic risk score (OR=1.14), and faecal calprotectin (OR=1.04). The same 3 variables were among the top 5 most important predictors as ranked by the random forest.\n\nConclusionA readily collectable panel of genetic risk variants, added to family history and faecal calprotectin, predicts those at greatest risk for developing CD with a good degree of accuracy.

epidemiology

Analysis of multistage in vitro fertilization data with mixed multilevel outcomes using joint modelling approaches

In vitro fertilization comprises a sequence of interventions concerned with the creation and culture of embryos which are then transferred to the patients uterus. While the clinically important endpoint is birth, the responses to each stage of treatment contain additional information about the reasons for success or failure. Joint analysis of the sequential responses is complicated by mixed outcome types defined at two levels (patient and embryo). We develop three methods for multistage analysis based on joining submodels for the different responses using latent variables and entering outcome variables as covariates for downstream responses. An application to routinely collected data is presented, and the strengths and limitations of each method are discussed.

epidemiology

Automating Mendelian randomization through machine learning to construct a putative causal map of the human phenome

A major application for genome-wide association studies (GWAS) has been the emerging field of causal inference using Mendelian randomization (MR), where the causal effect between a pair of traits can be estimated using only summary level data. MR depends on SNPs exhibiting vertical pleiotropy, where the SNP influences an outcome phenotype only through an exposure phenotype. Issues arise when this assumption is violated due to SNPs exhibiting horizontal pleiotropy. We demonstrate that across a range of pleiotropy models, instrument selection will be increasingly liable to selecting invalid instruments as GWAS sample sizes continue to grow. Methods have been developed in an attempt to protect MR from different patterns of horizontal pleiotropy, and here we have designed a mixture-of-experts machine learning framework (MR-MoE 1.0) that predicts the most appropriate model to use for any specific causal analysis, improving on both power and false discovery rates. Using the approach, we systematically estimated the causal effects amongst 2407 phenotypes. Almost 90% of causal estimates indicated some level of horizontal pleiotropy. The causal estimates are organised into a publicly available graph database (http://eve.mrbase.org), and we use it here to highlight the numerous challenges that remain in automated causal inference.

epidemiology

Multiscale Model Within-host and Between-host for Viral Infectious Diseases

Multiscale models possess the potential to uncover new insights into infectious diseases. Here, a rigorous stability analysis of a multiscale model within-host and between-host is presented. The within-host model describes virus replication and the respective immune response while disease transmission is represented by a simple susceptible-infected (SI) model.\n\nThe bridge of within-to between-host is by considering transmission as a function of the viral load of the within-host level. Consequently, stability and bifurcation analyses were developed coupling the two basic reproduction numbers [Formula] and [Formula]for the within- and the between-host subsystems, respectively. Local stability results for each subsystem, such as a unique stable equilibrium point, recapitulate classical approaches to infection and epidemic control.\n\nUsing a Lyapunov function, global stability of the between-host system was obtained. A main result was the derivation of the [Formula] as a general increasing function of [Formula]. Numerical analyses reveal that a Michaelis-Menten form based on the virus is more likely to recapitulate the behavior between the scales than a form directly proportional to the virus. Our work contributes basic understandings of the two models and casts light on the potential effects of the coupling function on linking the two scales.

epidemiology

Use of whole genome sequencing to investigate an outbreak of gonorrhoea among females in urban New South Wales, Australia, 2012 to 2014.

Increasing rates of gonorrhoea have been observed among urban heterosexuals within the Australian state of New South Wales (NSW). Here, we applied whole genome sequencing (WGS) to better understand transmission dynamics. Ninety-four isolates of a particular N. gonorrhoeae genotype (G122) associated with female patients (years 2012 to 2014) underwent phylogenetic analysis using core single nucleotide polymorphisms (SNPs). Context for genetic variation was provided by including an unbiased selection of 1,870 N. gonorrhoeae genomes from a recent United Kingdom (UK) study. NSW genomes formed a single clade, with the majority of isolates belonging to one of five clusters, and comprised patients of varying age groups. Intra-patient variability was less than 7 core SNPs. Several patients had indistinguishable core SNPs, suggesting a common infection source. These data have provided an enhanced understanding of transmission of N. gonorrhoeae among urban heterosexuals in NSW, Australia, and highlight the value of using WGS in N. gonorrhoeae outbreak investigations.

epidemiology

Assessing the performance of real-time epidemic forecasts

Real-time forecasts based on mathematical models can inform critical decision-making during infectious disease outbreaks. Yet, epidemic forecasts are rarely evaluated during or after the event, and there is little guidance on the best metrics for assessment. Here, we propose an evaluation approach that disentangles different components of forecasting ability using metrics that separately assess the calibration, sharpness and unbiasedness of forecasts. This makes it possible to assess not just how close a forecast was to reality but also how well uncertainty has been quantified. We used this approach to analyse the performance of weekly forecasts we generated in real time in Western Area, Sierra Leone, during the 2013-16 Ebola epidemic in West Africa. We investigated a range of forecast model variants based on the model fits generated at the time with a semi-mechanistic model, and found that good probabilistic calibration was achievable at short time horizons of one or two weeks ahead but models were increasingly inaccurate at longer forecasting horizons. This suggests that forecasts may have been of good enough quality to inform decision making requiring predictions a few weeks ahead of time but not longer, reflecting the high level of uncertainty in the processes driving the trajectory of the epidemic. Comparing forecasts based on the semi-mechanistic model to simpler null models showed that the best semi-mechanistic model variant performed better than the null models with respect to probabilistic calibration, and that this would have been identified from the earliest stages of the outbreak. As forecasts become a routine part of the toolkit in public health, standards for evaluation of performance will be important for assessing quality and improving credibility of mathematical models, and for elucidating difficulties and trade-offs when aiming to make the most useful and reliable forecasts.

epidemiology