Search bioRxivSearch

Biology subjects

De Anda, V.

Publications and source records attributed to De Anda, V..

2 recordsLinked to original sources

MEBS, a software platform to evaluate large (meta)genomic collections according to their metabolic machinery: unraveling the sulfur cycle

BACKGROUNDThe increasing number of metagenomic and genomic sequences has dramatically improved our understanding of microbial diversity, yet our ability to infer metabolic capabilities in such datasets remains challenging.\n\nFINDINGSWe describe the Multigenomic Entropy Based Score pipeline (MEBS), a software platform designed to evaluate, compare and infer complex metabolic pathways in large omic datasets, including entire biogeochemical cycles. MEBS is open source and available through https://github.com/eead-csic-compbio/metagenome_Pfam_score. To demonstrate its use we modeled the sulfur cycle by exhaustively curating the molecular and ecological elements involved (compounds, genes, metabolic pathways and microbial taxa). This information was reduced to a collection of 112 characteristic Pfam protein domains and a list of complete-sequenced sulfur genomes. Using the mathematical framework of relative entropy (H), we quantitatively measured the enrichment of these domains among sulfur genomes. The entropy of each domain was used to both: build up a final score that indicates whether a (meta)genomic sample contains the metabolic machinery of interest and to propose marker domains in metagenomic sequences such as DsrC (PF04358). MEBS was benchmarked with a dataset of 2,107 non-redundant microbial genomes from RefSeq and 935 metagenomes from MG-RAST. Its performance, reproducibility, and robustness were evaluated using several approaches, including random sampling, linear regression models, Receiver Operator Characteristic plots and the Area Under the Curve metric (AUC). Our results support the broad applicability of this algorithm to accurately classify (AUC=0.985) hard to culture genomes (e.g., Candidatus Desulforudis audaxviator), previously characterized ones and metagenomic environments such as hydrothermal vents, or deep-sea sediment.\n\nCONCLUSIONSOur benchmark indicates that an entropy-based score can capture the metabolic machinery of interest and be used to efficiently classify large genomic and metagenomic datasets, including uncultivated/unexplored taxa

bioinformatics

A new multi-genomic approach for the study of biogeochemical cycles at global scale: the molecular reconstruction of the sulfur cycle

Despite the great advances in microbial ecology and the explosion of high throughput sequencing, our ability to understand and integrate the global biogeochemical cycles is still limited. Here we propose a novel approach to summarize the complexity of the Sulfur cycle based on the minimum ecosystem concept, the microbial mat model and the relative entropy of protein domains involved in S-metabolism. This methodology produces a single value, called the Sulfur Score (SS), which informs about the specific S-related molecular machinery. After curating an inventory of microorganisms, pathways and genes taking part in this cycle, we benchmark the performance of the SS on a collection of 2,107 non-redundant RefSeq genomes, 900 metagenomes from MG-RAST and 35 metagenomes analyzed for the first time. We find that the SS is able to correctly classify microorganisms known to be involved in the S-cycle, yielding an Area Under the ROC Curve of 0.985. Moreover, when sorting environments the top-scoring metagenomes were hydrothermal vents, microbial mats and deep-sea sediments, among others. This methodology can be generalized to the analysis of other biogeochemical cycles or processes. Provided that an inventory of relevant pathways and microorganisms can be compiled, entropy-based scores could be used to detect environmental patterns and informative samples in multi-genomic scale.

bioinformatics