Search bioRxiv⌕ Search

Biology subjects

Ruiz, B.

Publications and source records attributed to Ruiz, B..

2 recordsLinked to original sources

SPARTA: Interpretable functional classification of microbiomes and detection of hidden cumulative effects.

The composition of the gut microbiota is a known factor in various diseases, and has proven to be a strong basis for automatic classification of disease state. A need for a better understanding of this community on the functional scale has since been voiced, as it would enhance these approaches biological interpretability. In this paper, we have developed a computational pipeline for integrating the functional annotation of the gut microbiota to an automatic classification process, and facilitating downstream interpretation of its results. The process takes as input taxonomic composition data (such as tables of Operational Taxonomic Unit (OTU) or Amplicon Sequence Variant (ASV) abundances), and links each component to its functional annotations through interrogation of the UniProt database. A functional profile of the gut microbiota is built from this basis. Both profiles, microbial and functional, are used to train Random Forest classifiers to discern unhealthy from control samples. An automatic variable selection is then performed on the basis of variable importance, and the method can be iterated until classification performances diminish. This process shows that the translation of the microbiota into functional profiles gives comparable, albeit slightly inferior performances when compared to microbial profiles. Through repetition, it also outputs a robust subset of discriminant variables. These selections were shown to be more reliable than those obtained by a state of the art method, and its contents were validated through a manual bibliographic research. The interconnections between selected OTUs and functional annotations were also analyzed, and revealed that important annotations emerge from the cumulated influence of non-selected OTUs.

bioinformatics↗

EsMeCaTa: Estimating metabolic capabilities from taxonomic affiliations

PurposeMetabarcoding, and metagenomic sequencing have enabled the characterization of highly diverse environmental communities. The challenge of estimating the metabolic functions carried out by these communities has led to the development of several state-of-the-art methods, most of which are tailored to a specific gene marker. However, the increasing diversity of approaches resulting from advances in sequencing technologies drives the need for methods capable of handling heterogeneous microbial community data. Moreover, predictions often depend on their internal analysis pipelines and are influenced by the underlying databases, which link marker genes to specific functional annotations. This limits users ability to evaluate the quality of predictions by tracing internal data and processes. Finally, users are constrained by the specific annotations provided by these methods (e.g. EC numbers), limiting their ability to conduct further specialized analyses based on intermediate results. MethodsEsMeCaTa predicts consensus proteomes and their associated functions from taxonomic affiliations. A key feature of EsMeCaTa is its explainability and flexibility. To support the flexible integration of heterogeneous sequencing data, EsMeCaTa utilizes taxonomic affiliations obtained through analyses of diverse sequencing datasets. To provide insight into the knowledge available for each taxonomic affliation and to interpret the relevance of predicted functions, EsMeCaTa identifies a taxonomic rank within a given affliation that is suffciently represented by documented proteomes in the UniProt database. The proteins of the UniProt proteomes are clustered and filtered according to a threshold to create consensus proteomes. These consensus proteomes are automatically annotated with functional information (e.g., EC numbers, GO terms) but they are also designed to be used in further customized annotation workflows. Functional annotations are reported in a functional table, which can be enriched with taxon abundances to generate comprehensive functional profiles. ResultsEsMeCaTa predictions have been validated using multiple datasets and compared to a state-of-the-art method. Additionally, it was applied to a novel metabarcoding dataset from a methanogenic reactor, characterizing the microbial community and biogas production across different time points and intake condition. Our results demonstrate the link between biogas production, intake condition and the dynamics of the metabolic functions predicted by EsMeCaTa in the microbial communities.

bioinformatics↗