Search bioRxiv⌕ Search

Biology subjects

Caron, A. R.

Publications and source records attributed to Caron, A. R..

4 recordsLinked to original sources

A general strategy for generating expert-guided, simplified views of ontologies

Annotation of biomedical entities with widely used, well-structured ontologies and ontology-aware tools ensures data and analyses are Findable, Accessible, Interoperable, and Reusable (FAIR). Standardized terms with synonyms support lexical search, while ontology structure enables biologically meaningful grouping of annotations, such as by location and type. However, ontologies serving diverse communities are often more complex than needed for specific applications, creating barriers to adoption by researchers and resource developers. For example, cell atlases often attempt simplifications by manually building term hierarchies linking to cell type and anatomy ontologies, but these may include relationship types unsuitable for grouping annotations. We present tools for validating human expert curated term hierarchies, developed in two human reference atlas projects, against ontology structures. The tools provide tabular statistics plus graphical views of matching and non-matching terms and relationships to support discussion and conflict resolution. The HuBMAP Human Reference Atlas (HRA) effort is used to validate the approach and tools, and the Human Developmental Cell Atlas is featured as a use case.

bioinformatics↗

The Unified Phenotype Ontology (uPheno): A framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

bioinformatics↗

The Ontology of Biological Attributes (OBA) - Computational Traits for the Life Sciences

Existing phenotype ontologies were originally developed to represent phenotypes that manifest as a character state in relation to a wild-type or other reference. However, these do not include the phenotypic trait or attribute categories required for the annotation of genome-wide association studies (GWAS), Quantitative Trait Loci (QTL) mappings or any population-focused measurable trait data. Moreover, variations in gene expression in response to environmental disturbances even without any genetic alterations can also be associated with particular biological attributes. The integration of trait and biological attribute information with an ever increasing body of chemical, environmental and biological data greatly facilitates computational analyses and it is also highly relevant to biomedical and clinical applications. The Ontology of Biological Attributes (OBA) is a formalised, species-independent collection of interoperable phenotypic trait categories that is intended to fulfil a data integration role. OBA is a standardised representational framework for observable attributes that are characteristics of biological entities, organisms, or parts of organisms. OBA has a modular design which provides several benefits for users and data integrators, including an automated and meaningful classification of trait terms computed on the basis of logical inferences drawn from domain-specific ontologies for cells, anatomical and other relevant entities. The logical axioms in OBA also provide a previously missing bridge that can computationally link Mendelian phenotypes with GWAS and quantitative traits. The term components in OBA provide semantic links and enable knowledge and data integration across specialised research community boundaries, thereby breaking silos.

bioinformatics↗

Specimen, Biological Structure, and Spatial Ontologies in Support of a Human Reference Atlas

The Human Reference Atlas (HRA) is defined as a comprehensive, three-dimensional (3D) atlas of all the cells in the healthy human body. It is compiled by an international team of experts that develop standard terminologies linked to 3D reference objects describing anatomical structures. The third HRA release (v1.2) covers spatial reference data and ontology annotations for 26 organs. Experts access the HRA annotations via spreadsheets and view reference models in 3D editing tools. This paper introduces the Common Coordinate Framework Ontology (CCFO) v2.0.1 that interlinks specimen, biological structure, and spatial data together with the CCF API which makes the HRA programmatically accessible and interoperable with Linked Open Data (LOD). We detail how real-world user needs and experimental data guide CCFO design and implementation, present CCFO classes and properties together with examples of their usage, and report on technical validation performed. The CCFO graph database and API are used in the HuBMAP portal, Virtual Reality Organ Gallery, and other applications that support data queries across multiple, heterogeneous sources.

bioinformatics↗