Search bioRxiv⌕ Search

Biology subjects

Bossy, R.

Publications and source records attributed to Bossy, R..

2 recordsLinked to original sources

MilkOligoCorpus: a semantically annotated resource for knowledge extraction on mammalian milk oligosaccharides

Milk oligosaccharides are bioactive components that regulate the composition of the neonatal microbiota and exert immunomodulatory functions. Their beneficial effects depend on their structure. Numerous studies have shown intra- and inter-species variation in the structural composition and concentration of these compounds in mammalian milk, yet the biological significance of such variation remains poorly understood. Automated natural language processing methods are promising tools for extracting and gathering structured data from unstructured texts to get insight into the biological significance of milk oligosaccharide variation across mammals. These methods require training and evaluation on manually annotated text corpora. While annotated corpora exist for chemical substances, none are specifically designed for training natural language processing models to extract information on milk oligosaccharides. To this end, we propose MilkOligoCorpus, a new gold standard for milk oligosaccharide composition in mammalian species. MilkOligoCorpus annotation scheme is a rich entity/relation model designed to describe the diversity pattern of milk oligosaccharides according to female factor variability and to help better understand the structure-related function of milk oligosaccharides. MilkOligoCorpus consists of 30 PubMed texts fully annotated with entities related to individuals, samples, oligosaccharides and oligosaccharide quantification linked by binary and n-ary relationships. To address data interoperability across disparate publications and databases, four terminological resources were also developed to assign unique identifiers to the entities, supported by external ontologies. This paper presents the creation of the MilkOligoCorpus and its associated schema, along with the development of annotation guidelines and terminological resources. We also present experimental results obtained by baseline information extraction models on the corpus.

bioinformatics↗

Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach

The dramatic increase in the amount of microbe descriptions in databases, reports and papers presents a two-fold challenge for accessing the information: integration of heterogeneous data in a standard ontology-based representation and normalization of the textual descriptions by semantic analysis. Recent text mining methods offer powerful ways to extract textual information and generate ontology-based representation. This paper describes the design of the Omnicrobe application that gathers comprehensive information on habitats, phenotypes and usages of microbes from scientific sources of high interest to the microbiology community. The Omnicrobe database contains around 1 million descriptions of microbe properties that are created by analyzing and combining six information sources of various kinds, i.e. biological resource catalogues, sequence database and scientific literature. The microbe properties are indexed by the Ontobiotope ontology and their taxa are indexed by an extended version of the taxonomy maintained by the National Center for Biotechnology Information. The Omnicrobe application covers all domains of microbiology. It provides an easy-to-use support in the resolution of scientific questions related to the habitats, phenotypes and uses of microbes through simple and complex ontology-based queries. We illustrate the potential of Omnicrobe with a use case from the food innovation domain.

bioinformatics↗