Search bioRxivSearch

Biology subjects

Vanni, C.

Publications and source records attributed to Vanni, C..

3 recordsLinked to original sources

AGNOSTOS-DB: a resource to unlock the uncharted regions of the coding sequence space

Genomes and metagenomes contain a considerable percentage of genes of unknown function, which are often excluded from downstream analyses limiting our understanding of the studied biological systems. To address this challenge, we developed AGNOSTOS, a combined database-computational workflow resource that unifies the known and unknown coding sequence space of genomes and metagenomes. Here, we present AGNOSTOS-DB, an extensive database of high-quality gene clusters enriched with functional, ecological and phylogenetic information. Moreover, AGNOSTOS allows integrating new data into existing AGNOSTOS-DBs, maximizing the information retrievable for the genes of unknown function. As a proof of concept, we provide a seed database that integrates the predicted genes from marine and human metagenomes, as well as from Bacteria, Archaea, Eukarya and giant viruses environmental and cultivar genomes. The seed database comprises 6,572,081 gene clusters connecting 342 million genes and represents a comprehensive and scalable resource for the inclusion and exploration of the unknown fraction of genomes and metagenomes.

microbiology

Functional repertoire convergence of distantly related eukaryotic plankton lineages revealed by genome-resolved metagenomics

Marine planktonic eukaryotes play a critical role in global biogeochemical cycles and climate. However, their poor representation in culture collections limits our understanding of the evolutionary history and genomic underpinnings of planktonic ecosystems. Here, we used 280 billion Tara Oceans metagenomic reads from polar, temperate, and tropical sunlit oceans to reconstruct and manually curate more than 700 abundant and widespread eukaryotic environmental genomes ranging from 10 Mbp to 1.3 Gbp. This genomic resource covers a wide range of poorly characterized eukaryotic lineages that complement long-standing contributions from culture collections while better representing plankton in the upper layer of the oceans. We performed the first comprehensive genome-wide functional classification of abundant unicellular eukaryotic plankton, revealing four major groups connecting distantly related lineages. Neither trophic modes of plankton nor its vertical evolutionary history could explain the functional repertoire convergence of major eukaryotic lineages that coexisted within oceanic currents for millions of years. CoverNavigating on the map of plankton genomics with Tara Oceans and anvio: a comprehensive genome-resolved metagenomic survey dedicated to eukaryotic plankton. O_FIG O_LINKSMALLFIG WIDTH=153 HEIGHT=200 SRC="FIGDIR/small/341214v2_ufig1.gif" ALT="Figure 1"> View larger version (82K): org.highwire.dtl.DTLVardef@536fe5org.highwire.dtl.DTLVardef@1d72cc9org.highwire.dtl.DTLVardef@1bd5281org.highwire.dtl.DTLVardef@739512_HPS_FORMAT_FIGEXP M_FIG C_FIG

microbiology

Light into the darkness: Unifying the known and unknown coding sequence space in microbiome analyses

Genes of unknown function are among the biggest challenges in molecular biology, especially in microbial systems, where 40%-60% of the predicted genes are unknown. Despite previous attempts, systematic approaches to include the unknown fraction into analytical workflows are still lacking. Here, we propose a conceptual framework and a computational workflow that bridge the known-unknown gap in genomes and metagenomes. We showcase our approach by exploring 415,971,742 genes predicted from 1,749 metagenomes and 28,941 bacterial and archaeal genomes. We quantify the extent of the unknown fraction, its diversity, and its relevance across multiple biomes. Furthermore, we provide a collection of 283,874 lineage-specific genes of unknown function for Cand. Patescibacteria, being a significant resource to expand our understanding of their unusual biology. Finally, by identifying a target gene of unknown function for antibiotic resistance, we demonstrate how we can enable the generation of hypotheses that can be used to augment experimental data.

microbiology