Search bioRxivSearch

Biology subjects

Quince, C.

Publications and source records attributed to Quince, C..

7 recordsLinked to original sources

Recent mixing of Vibrio parahaemolyticus populations

BackgroundHumans have profoundly affected the ocean environment but little is known about anthropogenic effects on the distribution of microbes. Vibrio parahaemolyticus is found in warm coastal waters and causes gastroenteritis in humans and economically significant disease in shrimps.\n\nResultsBased on data from 1,103 genomes, we show that V. parahaemolyticus is divided into four diverse populations, VppUS1, VppUS2, VppX and VppAsia. The first two are largely restricted to the US and Northern Europe, while the others are found worldwide, with VppAsia making up the great majority of isolates in the seas around Asia. Patterns of diversity within and between the populations are consistent with them having arisen by progressive divergence via genetic drift during geographical isolation. However, we find that there is substantial overlap in their current distribution. These observations can be reconciled without requiring genetic barriers to exchange between populations if dispersal between oceans has increased dramatically in the recent past. We found that VppAsia isolates from the US have an average of 1.01% more shared ancestry with VppUS1 and VppUS2 isolates than VppAsia isolates from Asia itself. Based on time calibrated trees of divergence within epidemic lineages, we estimate that recombination affects about 0.017% of the genome per year, implying that the genetic mixture has taken place within the last few decades.\n\nConclusionsThese results suggest that human activity, such as shipping and aquatic products trade, are responsible for the change of distribution pattern of this marine species.

microbiology

Systems genetic discovery of host-microbiome interactions reveals mechanisms of microbial involvement in disease

The role of the microbiome in health and disease involves complex networks of host genetics, genomics, microbes and environment. Identifying the mechanisms of these interactions has remained challenging. Systems genetics in the laboratory mouse enables data-driven discovery of network components and mechanisms of host-microbial interactions underlying multiple disease phenotypes. To examine the interplay among the whole host genome, transcriptome and microbiome, we mapped quantitative trait loci and correlated the abundance of cecal mRNA, luminal microflora, physiology and behavior in incipient strains of the highly diverse Collaborative Cross mouse population. The relationships that are extracted can be tested experimentally to ascribe causality among host and microbe in behavior and physiology, providing insight into disease. Application of this strategy in the Collaborative Cross population revealed experimentally validated mechanisms of microbial involvement in models of autism, inflammatory bowel disease and sleep disorder.\n\neTOC BlurbHost genetic diversity provides a variable selection environment and physiological context for microbiota and their interaction with host physiology. Using a highly diverse mouse population Bubier et al. identified a variety of host, microbe and potentially disease interactions.\n\nHighlights* 18 significant species-specific QTL regulating microbial abundance were identified\n* Cis and trans eQTL for 1,600 cecal transcripts were mapped in the Collaborative Cross\n* Sleep phenotypes were highly correlated with the abundance of B.P. Odoribacter\n* Elimination of sleep-associated microbes restored normal sleep patterns in mice.

genetics

Machine learning based prediction of functional capabilities in metagenomically assembled microbial genomes

The increasing popularity of genome resolved meta genomics - the binning of genomes of potentially uncultured organisms direct from the environmental DNA - has resulted in a deluge of draft genomes. There is a pressing need to develop methods to interpret this data. Here, we used machine learning to predict functional and metabolic traits of microbes from their genomes. We collated an extensive database of 84 phenotypic traits associated with 9407 prokaryotic genomes and trained different machine learning models on this data. We found that a lasso logistic regression based on the frequency of gene orthologs had the best combination of functional prediction performance and interpretability. This model was able to classify 65 phenotypic traits with greater than 90

microbiology

Accurate Reconstruction of Microbial Strains Using Representative Reference Genomes

Exploring the genetic diversity of microbes within the environment through metagenomic sequencing first requires classifying these reads into taxonomic groups. Current methods compare these sequencing data with existing biased and limited reference databases. Several recent evaluation studies demonstrate that current methods either lack sufficient sensitivity for species-level assignments or suffer from false positives, overestimating the number of species in the metagenome. Both are especially problematic for the identification of low-abundance microbial species, e. g. detecting pathogens in ancient metagenomic samples. We present a new method, SPARSE, which improves taxonomic assignments of metagenomic reads. SPARSE balances existing biased reference databases by grouping reference genomes into similarity-based hierarchical clusters, implemented as an efficient incremental data structure. SPARSE assigns reads to these clusters using a probabilistic model, which specifically penalizes non-specific mappings of reads from unknown sources and hence reduces false-positive assignments. Our evaluation on simulated datasets from two recent evaluation studies demonstrated the improved precision of SPARSE in comparison to other methods for species-level classification. In a third simulation, our method successfully differentiated multiple co-existing Escherichia coli strains from the same sample. In real archaeological datasets, SPARSE identified ancient pathogens with[≤] 0.02% abundance, consistent with published findings that required additional sequencing data. In these datasets, other methods either missed targeted pathogens or reported non-existent ones. SPARSE and all evaluation scripts are available at https://github.com/zheminzhou/SPARSE.

bioinformatics

Nitrogen-Fixing Populations Of Planctomycetes And Proteobacteria Are Abundant In The Surface Ocean

Nitrogen fixation in the surface ocean impacts the global climate by regulating the microbial primary productivity and the sequestration of carbon through the biological pump. Cyanobacterial populations have long been thought to represent the main suppliers of the bio-available nitrogen in this habitat. However, recent molecular surveys of nitrogenase reductase gene revealed the existence of rare non-cyanobacterial populations that can also fix nitrogen. Here, we characterize for the first time the genomic content of some of these heterotrophic bacterial diazotrophs (HBDs) inhabiting the open surface ocean waters. They represent new lineages within Planctomycetes and Proteobacteria, a phylum never linked to nitrogen fixation prior to this study. HBDs were surprisingly abundant in the Pacific Ocean and the Atlantic Ocean northwest, conflicting with decades of PCR surveys. The abundance and widespread occurrence of non-cyanobacterial diazotrophs in the surface ocean emphasizes the need to re-evaluate their role in the nitrogen cycle and primary productivity.

microbiology

Millennia of genomic stability within the invasive Para C Lineage of Salmonella enterica

Salmonella enterica serovar Paratyphi C is the causative agent of enteric (paratyphoid) fever. While today a potentially lethal infection of humans that occurs in Africa and Asia, early 20th century observations in Eastern Europe suggest it may once have had a wider-ranging impact on human societies. We recovered a draft Paratyphi C genome from the 800-year-old skeleton of a young woman in Trondheim, Norway, who likely died of enteric fever. Analysis of this genome against a new, significantly expanded database of related modern genomes demonstrated that Paratyphi C is descended from the ancestors of swine pathogens, serovars Choleraesuis and Typhisuis, together forming the Para C Lineage. Our results indicate that Paratyphi C has been a pathogen of humans for at least 1,000 years, and may have evolved after zoonotic transfer from swine during the Neolithic period.\n\nOne Sentence SummaryThe combination of an 800-year-old Salmonella enterica Paratyphi C genome with genomes from extant bacteria reshapes our understanding of this pathogens origins and evolution.

microbiology

Critical Assessment of Metagenome Interpretation - a benchmark of computational metagenomics software

In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.

bioinformatics