Search bioRxiv⌕ Search

Biology subjects

Blanquart, S.

Publications and source records attributed to Blanquart, S..

7 recordsLinked to original sources

FUSE-PhyloTree: Linking functions and sequence conservation modules of a protein family through phylogenomic analysis

FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g., paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the familys phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. Availability and ImplementationFUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree. Supplementary InformationAn illustration of the application of FUSE-PhyloTree to the fibulin protein family is presented in the Appendix.

bioinformatics↗

A duo of fungi and complex and dynamic bacterial community networks contribute to shape the Ascophyllum nodosum holobiont

The brown alga Ascophyllum nodosum and its microbiota form a dynamic functional entity named holobiont. Some microbial partners may play a role in seaweed health through bioactive compounds crucial for normal morphology, development, and physiological acclimation. However, the full spectrum of the microbial diversity and its variations according to algal life stage, season, and location have not been comprehensively studied. This study uses 208 short-read metabarcoding samples to characterize the bacterial, archaeal, and microeukaryotic communities of A. nodosum across three nearby sites, four thallus parts, and a monthly survey, aiming to explore the dynamics of ecological interactions within the holobiont. Our results revealed that A. nodosum harbors a predominantly bacterial microbiota, varying significantly across all covariables, while archaea were virtually absent. An innovative normalization using the co-amplified host reads provided an estimation of bacterial abundance, revealing a drastic decline in May, potentially linked to epidermal shedding. In contrast, fungal communities were stable, dominated by Mycophycias ascophylli and Moheitospora sp., which remained closely associated with the host year-round. We identified a core microbiome of 22 ASVs, consistently found in all samples, including Granulosicoccus, a genus consistently abundant in other brown algal microbiota. Sequence clustering revealed multiple species which vary according seasons, even in the overall stable Granulosicoccus genus. Co-occurrence network analysis revealed putative interactions between microbial groups in response to ecological niches. Overall, these findings highlight the dynamic of bacterial interactions and stable fungal associations within the A. nodosum holobiont, providing new insights into the ecology of its microbiota.

ecology↗

SPARTA: Interpretable functional classification of microbiomes and detection of hidden cumulative effects.

The composition of the gut microbiota is a known factor in various diseases, and has proven to be a strong basis for automatic classification of disease state. A need for a better understanding of this community on the functional scale has since been voiced, as it would enhance these approaches biological interpretability. In this paper, we have developed a computational pipeline for integrating the functional annotation of the gut microbiota to an automatic classification process, and facilitating downstream interpretation of its results. The process takes as input taxonomic composition data (such as tables of Operational Taxonomic Unit (OTU) or Amplicon Sequence Variant (ASV) abundances), and links each component to its functional annotations through interrogation of the UniProt database. A functional profile of the gut microbiota is built from this basis. Both profiles, microbial and functional, are used to train Random Forest classifiers to discern unhealthy from control samples. An automatic variable selection is then performed on the basis of variable importance, and the method can be iterated until classification performances diminish. This process shows that the translation of the microbiota into functional profiles gives comparable, albeit slightly inferior performances when compared to microbial profiles. Through repetition, it also outputs a robust subset of discriminant variables. These selections were shown to be more reliable than those obtained by a state of the art method, and its contents were validated through a manual bibliographic research. The interconnections between selected OTUs and functional annotations were also analyzed, and revealed that important annotations emerge from the cumulated influence of non-selected OTUs.

bioinformatics↗

Evolutionary genomics of the emergence of brown algae as key components of coastal ecosystems

Brown seaweeds are keystone species of coastal ecosystems, often forming extensive underwater forests, that are under considerable threat from climate change. Despite their ecological and evolutionary importance, this phylogenetic group, which is very distantly related to animals and land plants, is still poorly characterised at the genome level. Here we analyse 60 new genomes that include species from all the major brown algal orders. Comparative analysis of these genomes indicated the occurrence of several major events coinciding approximately with the emergence of the brown algal lineage. These included marked gain of new orthologous gene families, enhanced protein domain rearrangement, horizontal gene transfer events and the acquisition of novel signalling molecules and metabolic pathways. The latter include enzymes implicated in processes emblematic of the brown algae such as biosynthesis of the alginate-based extracellular matrix, and halogen and phlorotannin biosynthesis. These early genomic innovations enabled the adaptation of brown algae to their intertidal habitats. The subsequent diversification of the brown algal orders tended to involve loss of gene families, and genomic features were identified that correlated with the emergence of differences in life cycle strategy, flagellar structure and halogen metabolism. Analysis of microevolutionary patterns within the genus Ectocarpus indicated that deep gene flow between species may be an important factor in genome evolution on more recent timescales. Finally, we show that integration of large viral genomes has had a significant impact on brown algal genome content and propose that this process has persisted throughout the evolutionary history of the lineage.

genomics↗

AuCoMe: inferring and comparing metabolisms across heterogeneous sets of annotated genomes

Comparative analysis of Genome-Scale Metabolic Networks (GSMNs) may yield important information on the biology, evolution, and adaptation of species. However, it is impeded by the high heterogeneity of the quality and completeness of structural and functional genome annotations, which may bias the results of such comparisons. To address this issue, we developed AuCoMe - a pipeline to automatically reconstruct homogeneous GSMNs from a heterogeneous set of annotated genomes without discarding available manual annotations. We tested AuCoMe with three datasets, one bacterial, one fungal, and one algal, and demonstrated that it successfully reduces technical biases while capturing the metabolic specificities of each organism. Our results also point out shared metabolic traits and divergence points among evolutionarily distant species, such as algae, underlining the potential of AuCoMe to accelerate the broad exploration of metabolic evolution across the tree of life.

systems biology↗

EsMeCaTa: Estimating metabolic capabilities from taxonomic affiliations

PurposeMetabarcoding, and metagenomic sequencing have enabled the characterization of highly diverse environmental communities. The challenge of estimating the metabolic functions carried out by these communities has led to the development of several state-of-the-art methods, most of which are tailored to a specific gene marker. However, the increasing diversity of approaches resulting from advances in sequencing technologies drives the need for methods capable of handling heterogeneous microbial community data. Moreover, predictions often depend on their internal analysis pipelines and are influenced by the underlying databases, which link marker genes to specific functional annotations. This limits users ability to evaluate the quality of predictions by tracing internal data and processes. Finally, users are constrained by the specific annotations provided by these methods (e.g. EC numbers), limiting their ability to conduct further specialized analyses based on intermediate results. MethodsEsMeCaTa predicts consensus proteomes and their associated functions from taxonomic affiliations. A key feature of EsMeCaTa is its explainability and flexibility. To support the flexible integration of heterogeneous sequencing data, EsMeCaTa utilizes taxonomic affiliations obtained through analyses of diverse sequencing datasets. To provide insight into the knowledge available for each taxonomic affliation and to interpret the relevance of predicted functions, EsMeCaTa identifies a taxonomic rank within a given affliation that is suffciently represented by documented proteomes in the UniProt database. The proteins of the UniProt proteomes are clustered and filtered according to a threshold to create consensus proteomes. These consensus proteomes are automatically annotated with functional information (e.g., EC numbers, GO terms) but they are also designed to be used in further customized annotation workflows. Functional annotations are reported in a functional table, which can be enriched with taxon abundances to generate comprehensive functional profiles. ResultsEsMeCaTa predictions have been validated using multiple datasets and compared to a state-of-the-art method. Additionally, it was applied to a novel metabarcoding dataset from a methanogenic reactor, characterizing the microbial community and biogas production across different time points and intake condition. Our results demonstrate the link between biogas production, intake condition and the dynamics of the metabolic functions predicted by EsMeCaTa in the microbial communities.

bioinformatics↗

MATAM: Reconstruction Of Phylogenetic Marker Genes From Short Sequencing Reads In Metagenomes

MotivationAdvances in the sequencing of uncultured environmental samples, dubbed metagenomics, raise a growing need for accurate taxonomic assignment. Accurate identification of organisms present within a community is essential to understanding even the most elementary ecosystems. However, current high-throughput sequencing technologies generate short reads which partially cover full-length marker genes and this poses difficult bioinformatic challenges for taxonomy identification at high resolution\n\nResultsWe designed MATAM, a software dedicated to the fast and accurate targeted assembly of short reads sequenced from a genomic marker of interest. The method implements a stepwise process based on construction and analysis of a read overlap graph. It is applied to the assembly of 16S rRNA markers and is validated on simulated, synthetic and genuine metagenomes. We show that MATAM outperforms other available methods in terms of low error rates and recovered genome fractions and is suitable to provide improved assemblies for precise taxonomic assignments.\n\nAvailabilityhttps://github.com/bonsai-team/matam\n\nContactpierre.pericard@gmail.com, helene.touzet@univ-lille1.fr

bioinformatics↗