Search bioRxivSearch

Biology subjects

Bas E. Dutilh

Publications and source records attributed to Bas E. Dutilh.

4 recordsLinked to original sources

Contig annotation tool CAT robustly classifies assembled metagenomic contigs and long sequences

In modern-day metagenomics, there is an increasing need for robust taxonomic annotation of long DNA sequences from unknown micro-organisms. Long metagenomic sequences may be derived from assembly of short-read metagenomes, or from long-read single molecule sequencing. Here we introduce CAT, a pipeline for robust taxonomic classification of long DNA sequences. We show that CAT correctly classifies contigs at different taxonomic levels, even in simulated metagenomic datasets that are very distantly related from the sequences in the database. CAT is implemented in Python and the required scripts can be freely downloaded from Github.

Bioinformatics

Bottom-up ecology of the human microbiome: from metagenomes to metabolomes

The environmental metabolome is a dominant and essential factor shaping microbial communities. Thus, we hypothesized that metagenomic datasets could reveal the quantitative metabolic status of a given sample. Using a newly developed bottom-up ecology algorithm, we predicted high-resolution metabolomes of hundreds of metagenomic datasets from the human microbiome, revealing body-site specific metabolomes consistent with known metabolomics data, and suggesting that common cosmetics ingredients are some of the major metabolites shaping the human skin microbiome.

Bioinformatics

Ecogenomics and biogeochemical impacts of uncultivated globally abundant ocean viruses

Ocean microbes drive global-scale biogeochemical cycling1, but do so under constraints imposed by viruses on host community composition, metabolism, and evolutionary trajectories2-5. Due to sampling and cultivation challenges, genome-level viral diversity remains poorly described and grossly understudied in nature such that <1% of observed surface ocean viruses, even those that are abundant and ubiquitous, are known5. Here we analyze a global map of abundant, double stranded DNA (dsDNA) viruses and viral-encoded auxiliary metabolic genes (AMGs) with genomic and ecological contexts through the Global Ocean Viromes (GOV) dataset, which includes complete genomes and large genomic fragments from both surface and deep ocean viruses sampled during the Tara Oceans and Malaspina research expeditions6,7. A total of 15,222 epi- and mesopelagic viral populations were identified that comprised 867 viral clusters (VCs, approximately genus-level groups8,9). This roughly triples known ocean viral populations10, doubles known candidate bacterial and archaeal virus genera9, and near-completely samples epipelagic communities at both the population and VC level. Thirty-eight of the 867 VCs were identified as the most impactful dsDNA viral groups in the oceans, as these were locally or globally abundant and accounted together for nearly half of the viral populations in any GOV sample. Most of these were predicted in silico to infect dominant, ecologically relevant microbes, while two thirds of them represent newly described viruses that lacked any cultivated representative. Beyond these taxon-specific ecological observations, we identified 243 viral-encoded AMGs in GOV, only 95 of which were known. Deeper analyses of 4 of these AMGs revealed that abundant viruses directly manipulate sulfur and nitrogen cycling, and do so throughout the epipelagic ocean. Together these data provide a critically-needed organismal catalog and functional context to begin meaningfully integrating viruses into ecosystem models as key players in nutrient cycling and trophic networks.

Ecology

FORMAL: A model to identify organisms present in metagenomes using Monte Carlo Simulation

One of the major goals in metagenomics is to identify organisms present in the microbial community from a huge set of unknown DNA sequences. This profiling has valuable applications in multiple important areas of medical research such as disease diagnostics. Nevertheless, it is not a simple task, and many approaches that have been developed are slow and depend on the read length of the DNA sequences. Here we introduce an innovative and agile approach which k-mer and Monte Carlo simulation to profile and report abundant organisms present in metagenomic samples and their relative abundance without sequence length dependencies. The program was tested with a simulated metagenomes, and the results show that our approach predicts the organisms in microbial communities and their relative abundance.

Bioinformatics