Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Genetic variability in the promoter region of TNF-α gene in the reservoir of Junin virus, Calomys musculinus (Rodentia, Cricetidae)

In this study, we assessed the genetic variability of the promoter region of TNF- gene in three natural populations of the cricetid rodent Calomys musculinus. This species is the natural reservoir of Junin virus, the etiological agent of Argentine Hemorrhagic fever. We found different levels of variability and varying signatures of natural selection in populations with different epidemiological histories.

genetics

Genotype-phenotype association mining in bipolar disorder: market research meets complex genetics

Disentangling the etiology of common, complex diseases is a major challenge in genetic research. For bipolar disorder (BD), several genome-wide association studies (GWAS) have been performed. Similar to other complex disorders, major breakthroughs in explaining the high heritability of BD through GWAS have remained elusive. To overcome this dilemma, genetic research into BD, has embraced a variety of strategies such as the formation of large consortia to increase sample size and sequencing approaches. Here we advocate a complementary approach making use of already existing GWAS data: applying a data mining procedure to identify yet undetected genotype-phenotype relationships. We adapted association rule mining, a data mining technique traditionally used in retail market research, to identify frequent and characteristic genotype patterns showing strong associations to phenotype clusters. We applied this strategy to three independent GWAS datasets from 2,835 phenotypically characterized patients with BD. In a discovery step, 20,882 candidate association rules were extracted. Two of these - one associated with eating disorder and the other with anxiety - remained significant in an independent dataset after robust correction for multiple testing, showing considerable effect sizes (odds ratio ~ 3.4 and 3.0, respectively). Our approach may help detect novel specific genotype-phenotype relationships in BD typically not explored by analyses like GWAS. While we adapted the data mining tool within the context of BD gene discovery, it may facilitate identifying highly specific genotype-phenotype relationships in subsets of genome-wide data sets of other complex phenotype with similar epidemiological properties and challenges to gene discovery efforts.

genomics

Phandango: an interactive viewer for bacterial population genomics.

SummaryFully exploiting the wealth of data in current bacterial population genomics datasets requires synthesising and integrating different types of analysis across millions of base pairs in hundreds or thousands of isolates. Current approaches often use static representations of phylogenetic, epidemiological, statistical and evolutionary analysis results that are difficult to relate to one another. Phandango is an interactive application running in a web browser allowing fast exploration of large-scale population genomics datasets combining the output from multiple genomic analysis methods in an intuitive and interactive manner.\n\nAvailabilityPhandango is a web application freely available for use at https://jameshadfield.github.io/phandango and includes a diverse collection of datasets as examples. Source code together with a detailed wiki page is available on GitHub at https://github.com/jameshadfield/phandango\n\nContactjh22@sanger.ac.uk, sh16@sanger.ac.uk

bioinformatics

Plasmodium simium causing human malaria: a zoonosis with outbreak potential in the Rio de Janeiro Brazilian Atlantic forest

BackgroundMalaria was eliminated from Southern and Southeastern Brazil over 50 years ago. However, an increasing number of autochthonous episodes attributed to Plasmodium vivax have been recently reported in the Atlantic forest region of Rio de Janeiro State. As P. vivax-like non-human primate malaria parasite species Plasmodium simium is locally enzootic, we performed a molecular epidemiological investigation in order to determine whether zoonotic malaria transmission is occurring.\n\nMethodsBlood samples of humans presenting signs and/or symptoms suggestive of malaria as well as from local howler-monkeys were examined by microscopy and PCR. Additionally, a molecular assay based on sequencing of the parasite mitochondrial genome was developed to distinguish between P. vivax and P. simium, and applied to 33 cases from outbreaks occurred in 2015 and 2016.\n\nResultsOf 28 samples for which the assay was successfully performed, all were shown to be P. simium, indicating the zoonotic transmission of this species to humans in this region. Sequencing of the whole mitochondrial genome of three of these cases showed that P. simium is most closely related to P. vivax parasites from South American.\n\nFindingsThe explored malaria outbreaks were caused by P. simium, previously considered a monkey-specific malaria parasite, related to but distinct from P. vivax, and which has never conclusively been shown to infect humans before.\n\nInterpretationThis unequivocal demonstration of zoonotic transmission, 50 years after the only previous report of P. simium in man, leads to the possibility that this parasite has always infected humans in this region, but that it has been consistently misdiagnosed as P. vivax due to a lack of molecular typing techniques. Thorough screening of the local non-human primate and anophelines is required to evaluate the extent of this newly recognized zoonotic threat to public health and malaria eradication in Brazil.\n\nFundingFundacao Carlos Chagas Filho de Amparo a Pesquisa do Estado de Rio de Janeiro (Faperj), The Brazilian National Council for Scientific and Technological Development (CNPq), JSPS Grant-in-Aid for scientific research, Secretary for Health Surveillance (SVS) of the Ministry of Health, Global Fund, and PRONEX Program of the CNPq.

genetics

Methicillin resistant Staphylococcus aureus emerged long before the introduction of methicillin in to clinical practice

The spread of drug-resistant bacterial pathogens pose a major threat to global health. It is widely recognised that the widespread use of antibiotics has generated selective pressures that have driven the emergence of resistant strains. Methicillin-resistant Staphylococcus aureus (MRSA) was first observed in 1960, less than one year after the introduction of this second generation {beta}-lactam antibiotic into clinical practice. Epidemiological evidence has always suggested that resistance arose around this period, when the mecA gene encoding methicillin resistance carried on an SCCmec element, was horizontally transferred to an intrinsically sensitive strain of S. aureus. Whole genome sequencing a collection of the very first MRSA isolates allowed us to reconstruct the evolutionary history of the archetypal MRSA. Bayesian phylogenetic reconstruction was applied to infer the time point at which this early MRSA lineage arose and when SCCmec was acquired. MRSA emerged in the mid 1940s, following the acquisition of an ancestral type I SCCmec element, some fourteen years prior to the first therapeutic use of methicillin. Methicillin use was not the original driving factor in the evolution of MRSA as previously thought. Rather it was the widespread use of first generation {beta}-lactams such as penicillin in the years prior to the introduction of methicillin, which selected for S. aureus strains carrying the mecA determinant. Crucially this highlights how new drugs, introduced to circumvent known resistance mechanisms, can be rendered ineffective by unrecognised adaptations in the bacterial population due to the historic selective landscape created by the widespread use of other antibiotics.

microbiology

Causal Analyses, Statistical Efficiency And Phenotypic Precision Through Recall-By-Genotype Study Design

Genome-wide association studies have been useful in identifying common genetic variants related to a variety of complex traits and diseases; however, they are often limited in their ability to inform about underlying biology. Whilst bioinformatics analyses, studies of cells, animal models and applied genetic epidemiology have provided some understanding of genetic associations or causal pathways, there is a need for new genetic studies that elucidate causal relationships and mechanisms in a cost-effective, precise and statistically efficient fashion. We discuss the motivation for and the characteristics of the Recall-by-Genotype (RbG) study design, an approach that enables genotype-directed deep-phenotyping and improvement in drawing causal inferences. Specifically, we present RbG designs using single and multiple variants and discuss the inferential properties, analytical approaches and applications of both. We consider the efficiency of the RbG approach, the likely value of RbG studies for the causal investigation of disease aetiology and the practicalities of incorporating genotypic data into population studies in the context of the RbG study design. Finally, we provide a catalogue of the UK-based resources for such studies, an online tool to aid the design of new RbG studies and discuss future developments of this approach.

genetics

Meffil: efficient normalisation and analysis of very large DNA methylation samples

AbstractO_ST_ABSBackgroundC_ST_ABSTechnological advances in high throughput DNA methylation microarrays have allowed dramatic growth of a new branch of epigenetic epidemiology. DNA methylation datasets are growing ever larger in terms of the number of samples profiled, the extent of genome coverage, and the number of studies being meta-analysed. Novel computational solutions are required to efficiently handle these data.\n\nMethodsWe have developed meffil, an R package designed to quality control, normalize and perform epigenome-wide association studies (EWAS) efficiently on large samples of Illumina Infinium HumanMethylation450 and MethylationEPIC BeadChip microarrays. We tested meffil by applying it to 6000 450k microarrays generated from blood collected for two different datasets, Accessible Resource for Integrative Epigenomic Studies (ARIES) and The Genetics of Overweight Young Adults (GOYA) study.\n\nResultsA complete reimplementation of functional normalization minimizes computational memory requirements to 5% of that required by other R packages, without increasing running time. Incorporating fixed and random effects alongside functional normalization, and automated estimation of functional normalisation parameters reduces technical variation in DNA methylation levels, thus reducing false positive associations and improving power. We also demonstrate that the ability to normalize datasets distributed across physically different locations without sharing any biologically-based individual-level data may reduce heterogeneity in meta-analyses of epigenome-wide association studies. However, we show that when batch is perfectly confounded with cases and controls functional normalization is unable to prevent spurious associations.\n\nConclusionsmeffil is available online (https://github.com/perishky/meffil/) along with tutorials covering typical use cases.

bioinformatics

Bank Vole Immunoheterogeneity May Limit Nephropatia Epidemica Emergence In A French Non-Endemic Region

Ecoevolutionary processes affecting hosts, vectors and pathogens are important drivers of zoonotic disease emergence. In this study, we focused on nephropathia epidemica (NE), which is caused by Puumala hantavirus (PUUV) whose natural reservoir is the bank vole, Myodes glareolus. Despite the continuous distribution of the reservoir in Europe, PUUV occurence is highly fragmented. We questioned the possibility of NE emergence in a French region that is considered to be NE-free but that is adjacent to a NE-endemic region. We first confirmed the epidemiology of these two regions using serological and virological surveys. We used bank vole population genetics to demonstrate the absence of spatial barriers that could have limited dispersal, and consequently, the spread of PUUV into the NE-free region. We next tested whether regional immunoheterogeneity could impact PUUV chances to establish, circulate and persist in the NE-free region. Immune responsiveness was phenotyped both in the wild and during experimental infections, using serological, virological and immune related gene expression assays. We showed that bank voles from the NE-free region were sensitive to experimental PUUV infection. We observed high levels of immunoheterogeneity between individuals and also between regions. In natural populations, antiviral gene expression (Tnf and Mx2 genes) reached higher levels in bank voles from the NE-free region. During experimental infections, anti-PUUV antibody production was higher in bank voles from the NE endemic region. Altogether, our results indicated a lower susceptibility to PUUV for bank voles from this NE-free region, what might limit PUUV circulation and persistence, and in turn, the risk of NE.\n\nGraphical abstract\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=128 SRC=\"FIGDIR/small/130252_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (25K):\norg.highwire.dtl.DTLVardef@f90a69org.highwire.dtl.DTLVardef@1a9e0dorg.highwire.dtl.DTLVardef@17e96b8org.highwire.dtl.DTLVardef@1d936e8_HPS_FORMAT_FIGEXP M_FIG C_FIG

ecology

Image Processing and Quality Control for the first 10,000 Brain Imaging Datasets from UK Biobank

UK Biobank is a large-scale prospective epidemiological study with all data accessible to researchers worldwide. It is currently in the process of bringing back 100,000 of the original participants for brain, heart and body MRI, carotid ultrasound and low-dose bone/fat x-ray. The brain imaging component covers 6 modalities (T1, T2 FLAIR, susceptibility weighted MRI, Resting fMRI, Task fMRI and Diffusion MRI). Raw and processed data from the first 10,000 imaged subjects has recently been released for general research access. To help convert this data into useful summary information we have developed an automated processing and QC (Quality Control) pipeline that is available for use by other researchers. In this paper we describe the pipeline in detail, following a brief overview of UK Biobank brain imaging and the acquisition protocol. We also describe several quantitative investigations carried out as part of the development of both the imaging protocol and the processing pipeline.

neuroscience

Population Genomics Of Cryptococcus neoformans var. grubii Reveals New Biogeographic Relationships And Finely Maps Hybridization

Cryptococcus neoformans var. grubii is the causative agent of cryptococcal meningitis, a significant source of mortality in immunocompromised individuals, typically HIV/AIDS patients from developing countries. Despite the worldwide emergence of this ubiquitous infection, little is known about the global molecular epidemiology of this fungal pathogen. Here we sequence the genomes of 188 diverse isolates and characterized the major subdivisions, their relative diversity and the level of genetic exchange between them. While most isolates of C. neoformans var. grubii belong to one of three major lineages (VNI, VNII, and VNB), some haploid isolates show hybrid ancestry including some that appear to have recently interbred, based on the detection of large blocks of each ancestry across each chromosome. Many isolates display evidence of aneuploidy, which was detected for all chromosomes. In diploid isolates of C. neoformans var. grubii (serotype A/A) and of hybrids with C. neoformans var. neoformans (serotype A/D) such aneuploidies have resulted in loss of heterozygosity, where a chromosomal region is represented by the genotype of only one parental isolate. Phylogenetic and population genomic analyses of isolates from Brazil revealed that the previously African VNB lineage occurs naturally in the South American environment. This suggests migration of the VNB lineage between Africa and South America prior to its diversification, supported by finding ancestral recombination events between isolates from different lineages and regions. The results provide evidence of substantial population structure, with all lineages showing multi-continental distributions demonstrating the highly dispersive nature of this pathogen.\n\nAuthor SummaryCryptococcus neoformans var. grubii is a human fungal pathogen of immunocompromised individuals that has global clinical impact, causing half a million deaths per year. Substantial genetic substructure exists for this pathogen, with two lineages found globally (VNI, VNII) whereas a third has appeared confined to sub-Saharan Africa (VNB). Here, we utilized genome sequencing of a large set of global isolates to examine the genetic diversity, hybridization, and biogeography of these lineages. We found that while the three major lineages are well separated, recombination between the lineages has occurred, notably resulting in hybrid isolates with segmented ancestry across the genome. In addition, we showed that isolates from South America are placed within the VNB lineage, formerly thought to be confined to Africa, and that there is phylogenetic separation between these geographies that substantially expands the diversity of these lineages. Our findings provide a new framework for further studies of the dynamics of natural populations of C. neoformans var. grubii.

genomics

A Supervised Statistical Learning Approach For Accurate Legionella pneumophila Source Attribution During Outbreaks

Public health agencies are increasingly relying on genomics during Legionnaires disease investigations. However, the causative bacterium (Legionella pneumophila) has an unusual population structure with extreme temporal and spatial genome sequence conservation. Furthermore, Legionnaires disease outbreaks can be caused by multiple L. pneumophila genotypes in a single source. These factors can confound cluster identification using standard phylogenomic methods. Here, we show that a statistical learning approach based on\n\nL. pneumophila core genome single nucleotide polymorphism (SNP) comparisons eliminates ambiguity for defining outbreak clusters and accurately predicts exposure sources for clinical cases. We illustrate the performance of our method by genome comparisons of 234 L. pneumophila isolates obtained from patients and cooling towers in Melbourne, Australia between 1994 and 2014. This collection included one of the largest reported Legionnaires disease outbreaks, involving 125 cases at an aquarium. Using only sequence data from L. pneumophila cooling tower isolates and including all core genome variation, we built a multivariate model using discriminant analysis of principal components (DAPC) to find cooling tower-specific genomic signatures, and then used it to predict the origin of clinical isolates. Model assignments were 93% congruent with epidemiological data, including the aquarium Legionnaires outbreak and three other unrelated outbreak investigations. We applied the same approach to a recently described investigation of Legionnaires disease within a UK hospital and observed model predictive ability of 86%. We have developed a promising means to breach L. pneumophila genetic diversity extremes and provide objective source attribution data for outbreak investigations.

microbiology

Shared Genetic Architecture Of Asthma With Allergic Diseases: A Genome-wide Cross Trait Analysis Of 112,000 Individuals From UK Biobank

Clinical and epidemiological data suggest that asthma and allergic diseases are associated. And may share a common genetic etiology. We analyzed genome-wide single-nucleotide polymorphism (SNP) data for asthma and allergic diseases in 35,783 cases and 76,768 controls of European ancestry from the UK Biobank. Two publicly available independent genome wide association studies (GWAS) were used for replication. We have found a strong genome-wide genetic correlation between asthma and allergic diseases (rg = 0.75, P = 6.84x10-62). Cross trait analysis identified 38 genome-wide significant loci, including novel loci such as D2HGDH and GAL2ST2. Computational analysis showed that shared genetic loci are enriched in immune/inflammatory systems and tissues with epithelium cells. Our work identifies common genetic architectures shared between asthma and allergy and will help to advance our understanding of the molecular mechanisms underlying co-morbid asthma and allergic diseases.

genetics

Salmonella enterica Serovar Typhimurium ST313 Responsible For Gastroenteritis In The UK Are Genetically Distinct From Isolates Causing Bloodstream Infections In Africa

The ST313 sequence type of Salmonella enterica serovar Typhimurium causes invasive non-typhoidal salmonellosis amongst immunocompromised people in sub-Saharan Africa (sSA). Previously, two distinct phylogenetic lineages of ST313 have been described which have rarely been found outside sSA. Following the introduction of routine whole genome sequencing of Salmonella enterica by Public Health England in 2014, we have discovered that 2.7% (79/2888) of S. Typhimurium from patients in England and Wales are ST313. Of these isolates, 59/72 originated from stool and 13/72 were from extra-intestinal sites. The isolation of ST313 from extra-intestinal sites was significantly associated with travel to Africa (OR 12 [95% CI: 3,53]). Phylogenetic analysis revealed previously unsampled diversity of ST313, and distinguished UK-linked isolates causing gastroenteritis from African-associated isolates causing invasive disease. Bayesian evolutionary investigation suggested that the two African lineages diverged from their most recent common ancestors independently, circa 1796 and 1903. The majority of genome degradation of African ST313 lineage 2 is conserved in the UK ST313 lineages and only 10/44 pseudogenes were lineage 2-specific. The African lineages carried a characteristic prophage and antibiotic resistance gene repertoire, suggesting a strong selection pressure for these horizontally-acquired genetic elements in the sSA setting. We identified an ST313 isolate associated with travel to Kenya that carried a chromosomally-located blaCTX-M-15, demonstrating the continual evolution of this sequence type in Africa in response to selection pressure exerted by antibiotic usage.\n\nThe S. Typhimurium ST313 sequence type has been primarily associated with invasive disease in Africa. Here, we highlight the power of routine whole-genome-sequencing by public health agencies to make epidemiologically-significant deductions that would be missed by conventional microbiological methods. The discovery of ST313 isolates responsible for gastroenteritis in the UK reveals new diversity in this important sequence type. We speculate that the niche specialization of sub-Saharan African ST313 lineages is driven in part by the acquisition of accessory genome elements.

microbiology

Curve Selection For Predicting Breast Cancer Metastasis From Prospective Gene Expression In Blood

We investigate whether there is information in gene expression levels in blood that predicts breast cancer metastasis. Our data comes from the NOWAC epidemiological cohort study where blood samples were provided at enrollment. This could be anywhere from years to weeks before any cancer diagnosis. When and if a cancer is diagnosed, it could be so in different ways: at a screening, between screenings, or in the clinic, outside of the screening program. To build predictive models we propose that variable selection should include followup time and stratify by detection method. We show by simulations that this improves the probability of selecting relevant predictor genes. We also demonstrate that it leads to improved predictions and more stable gene signatures in our data. There is some indication that blood gene expression levels hold predictive information about metastasis. With further development such information could be used for early detection of metastatic potential and as such aid in cancer treatment.

bioinformatics

Regional Polygenic Covariance Reveals Heterogeneity In The Shared Heritability Between Complex Traits

BackgroundComplex traits can share a substantial proportion of their polygenic heritability. However, genome-wide polygenic correlations between pairs of traits can mask heterogeneity in their shared polygenic effects across loci. We propose a novel method (WML-RPC) to evaluate polygenic correlation between two complex traits in small genomic regions using summary association statistics. Our method tests for evidence that the polygenic effect at a given region affects two traits concurrently.\n\nResultsWe show through simulations that our method is well calibrated, powerful and more robust to misspecification of linkage disequilibrium than other methods under a polygenic model. As small genomic regions are more likely to harbour specific genetic effects, our method is ideal to identify heterogeneity in shared polygenic correlation across regions. We illustrate the usefulness of our method by addressing two questions related to cardio-metabolic traits. First, we explored how regional polygenic correlation can inform on the strong epidemiological association between HDL cholesterol and coronary artery disease (CAD), suggesting a key role for triglycerides metabolism. Second, we investigated the potential role of PPAR{gamma} activators in the prevention of CAD.\n\nConclusionsOur results provide a compelling argument that shared heritability between complex traits is highly heterogeneous across loci.

genetics

Effects of Prenatal Stress on Structural Brain Development and Aging in Humans

Healthy brain aging is a major determinant of quality of life, allowing integration into society at all ages. Human epidemiological and animal studies indicate that in addition to lifestyle and genetic factors, environmental influences in prenatal life have a major impact on brain aging and age-associated brain disorders. The aim of this review is to summarize the existing literature on the consequences of maternal anxiety, stress, and malnutrition for structural brain aging and predisposition for age-associated brain diseases, focusing on studies with human samples. In conclusion, the results underscore the importance of a healthy mother-child relationship, starting in pregnancy, and the need for early interventions if this relationship is compromised.

neuroscience

The mismeasure of (obese) man - body mass index is an inadequate indicator of body condition

The body mass index (BMI), recommended also by the World Health Organization, is currently used as the leading body condition indicator in clinical and epidemiological studies and has become popular among the general public. Here we provide evidence of a systematic bias in BMI, showing that BMI is dependent on body height. As a result, shorter persons have a greater chance of being classified as underweight, while taller persons as overweight, even if they have identical nutritional status. Use of BMI should be thus abandoned in diagnosis as well as in clinical and experimental studies.

physiology

The effects of mutational process and selection on driver mutations across cancer types

Epidemiological evidence has long associated environmental mutagens with increased cancer risk. However, links between specific mutation-causing processes and the acquisition of individual driver mutations have remained obscure. Here we have used public cancer sequencing data to infer the independent effects of mutation and selection on driver mutation complement. First, we detect associations between a range of mutational processes, including those linked to smoking, ageing, APOBEC and DNA mismatch repair (MMR) and the presence of key driver mutations across cancer types. Second, we quantify differential selection between well-known alternative driver mutations, including differences in selection between distinct mutant residues in the same gene. These results show that while mutational processes play a large role in determining which driver mutations are present in a cancer, the role of selection frequently dominates.

cancer biology