Search bioRxiv⌕ Search

Biology subjects

Pava, D.

Publications and source records attributed to Pava, D..

2 recordsLinked to original sources

AI semantics for biomedical data integration

Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge. Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries.

bioinformatics↗

International Mouse Phenotyping Consortium: Investigating gene function and providing insights into human disease.

The International Mouse Phenotyping Consortium (IMPC; https://www.mousephenotype.org/) web portal contains phenotype data for mouse protein-coding genes derived from analysis of data obtained in a systematic and high-throughput fashion from knock-out lines produced by IMPC. The project has produced >1,400 mouse models of human disease that recapitulate phenotypes observed in patients. Over 8000 papers rely on data or reagents generated by IMPC, demonstrating the impact of the project on the research and clinical communities, and IMPC data is incorporated into other resources, such as MGI, Open Targets and UniProt. Data release (DR23.0, 2025) contains > 100 million data points from 9,277 genes and identified 113,803 significant phenotypes. To manage efficient access to this quantity of high dimensional data the IMPC web portal has been rebuilt using a cloud native architecture. The modern user interface retains the look and feel of the original portal with improvements identified through a usability study. New data visualisation and training materials for large scale data access through the API have also been developed to make the resource easier to use. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/681205v1_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@199243borg.highwire.dtl.DTLVardef@119c12forg.highwire.dtl.DTLVardef@1da1a27org.highwire.dtl.DTLVardef@1eb141d_HPS_FORMAT_FIGEXP M_FIG C_FIG

genetics↗