Search bioRxiv⌕ Search

Biology subjects

Le Guillarme, N.

Publications and source records attributed to Le Guillarme, N..

3 recordsLinked to original sources

A Practical Approach to Constructing a Knowledge Graph for Soil Ecological Research

With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still poorly equipped when it comes to integrating disparate datasets into a unified knowledge graph with well-defined semantics. This paper presents a practical approach to constructing a knowledge graph from heterogeneous and distributed (semi-)structured data sources. To illustrate our approach, we integrate several datasets on the trophic ecology of soil organisms into a trophic knowledge graph and show how information can be retrieved from the graph to support multi-trophic studies.

ecology↗

The Soil Food Web Ontology: aligning trophic groups, processes, and resources to harmonise and automatise soil food web reconstructions

Although soil ecology has benefited from recent advances in describing the functional and trophic traits of soil organisms, data reuse for large-scale soil food-web reconstructions still faces challenges. These obstacles include: (1) most data on the trophic interactions and feeding behaviour of soil organisms being scattered across disparate repositories, without well-established standard for describing and structuring trophic datasets; (2) the existence of various competing terms, rather than consensus, to delineate feeding-related concepts such as diets, trophic groups, feeding processes, resource types, leading to ambiguities that hinder meaningful data integration from different studies; (3) considerable divergence in the trophic classification of numerous soil organisms, or even the lack of such classifications, leading to discrepancies in the resolution of reconstructed food webs and complicating the reuse and comparison of food-web models within synthetic studies. To address these issues, we introduce the Soil Food Web Ontology, a novel formal conceptual framework designed to foster agreement on the trophic ecology of soil organisms. This ontology represents a collaborative and ongoing endeavour aimed at establishing consensus and formal definitions for the array of concepts relevant to soil trophic ecology. Its primary objective is to enhance the accessibility, interpretation, combination, reuse, and automated processing of trophic data. By harmonising the terminology and fundamental principles of soil trophic ecology, we anticipate that the Soil Food Web Ontology will improve knowledge management within the field. It will help soil ecologists to better harness existing information regarding the feeding behaviours of soil organisms, facilitate more robust trophic classifications, streamline the reconstruction of soil food webs, and ultimately render food-web research more inclusive, reusable and reproducible.

ecology↗

TaxoNERD: deep neural models for the recognition of taxonomic entities in the ecological and evolutionary literature

O_LIGiven the biodiversity crisis, we more than ever need to access information on multiple taxa (e.g. distribution, traits, diet) in the scientific literature to understand, map and predict all-inclusive biodiversity. Tools are needed to automatically extract useful information from the ever-growing corpus of ecological texts and feed this information to open data repositories. A prerequisite is the ability to recognise mentions of taxa in text, a special case of named entity recognition (NER). In recent years, deep learning-based NER systems have become ubiquitous, yielding state-of-the-art results in the general and biomedical domains. However, no such tool is available to ecologists wishing to extract information from the biodiversity literature. C_LIO_LIWe propose a new tool called TaxoNERD that provides two deep neural network (DNN) models to recognise taxon mentions in ecological documents. To achieve high performance, DNN-based NER models usually need to be trained on a large corpus of manually annotated text. Creating such a gold standard corpus (GSC) is a laborious and costly process, with the result that GSCs in the ecological domain tend to be too small to learn an accurate DNN model from scratch. To address this issue, we leverage existing DNN models pretrained on large biomedical corpora using transfer learning. The performance of our models is evaluated on four GSCs and compared to the most popular taxonomic NER tools. C_LIO_LIOur experiments suggest that existing taxonomic NER tools are not suited to the extraction of ecological information from text as they performed poorly on ecologically-oriented corpora, either because they do not take account of the variability of taxon naming practices, or because they do not generalise well to the ecological domain. Conversely, a domain-specific DNN-based tool like TaxoNERD outperformed the other approaches on an ecological information extraction task. C_LIO_LIEfforts are needed in order to raise ecological information extraction to the same level of performance as its biomedical counterpart. One promising direction is to leverage the huge corpus of unlabelled ecological texts to learn a language representation model that could benefit downstream tasks. These efforts could be highly beneficial to ecologists on the long term. C_LI

bioinformatics↗