Search bioRxiv⌕ Search

Biology subjects

Yeh, F.-Y.

Publications and source records attributed to Yeh, F.-Y..

3 recordsLinked to original sources

SEA CDM: Study-Experiment-Assay Common Data Model and Databases for Cross-Domain Data Integration and Analysis

With the increasing volume of biomedical experimental data, standardizing, sharing, and integrating heterogeneous experimental data across domains has become a major challenge. To address this challenge, we have developed an ontology-supported Study-Experiment-Assay (SEA) common data model (CDM), which includes 10 core and 3 auxiliary classes based on object-oriented modeling. SEA CDM uses interoperable ontologies for data standardization and knowledge inference. Building on the SEA CDM, we developed the Ontology-based SEA Network (OSEAN) relational database and knowledge graph, along with a set of ETL (Extract, Transform, Load) and query tools, and further applied them to represent 1,278 immune studies with over two million samples from three resources: VIGET, ImmPort, and CELLxGENE. Using simple, robust queries and analyses, our research identified multiple scientific insights into sex-specific immune responses, such as neutrophil degranulation and TNF binding to physiological receptors, following live attenuated and trivalent inactivated influenza vaccination. The novel SEA CDM system lays a foundation for establishing an integrative biodata ecosystem across biological and biomedical domains.

bioinformatics↗

VO: The Vaccine Ontology

With the widespread use of vaccines in research and clinical settings, there is an urgent need to standardize vaccine representation, integrate information across diverse vaccine types, and support computer-assisted reasoning. Accordingly, we have since 2007 developed the community-based Vaccine Ontology (VO), which aligns with the Basic Formal Ontology and adheres to OBO Foundry principles. VO models ontologically vaccines, vaccine components, vaccine immune responses, vaccine investigation studies and other vaccine-related topics. VO represents more than 10,000 vaccines targeting 289 infectious pathogens and cancers in humans and over 30 nonhuman animal species. VO provides mappings to external resources such as RxNorm, CVX, FDA, and USDA. Various VO use cases exist. VO facilitates vaccine standardization in resources such as the VIOLIN vaccine database, ImmPort, and the Vaccine Adjuvant Compendium (VAC). Semantic queries can be made to query VO. VO has been shown to enhance experimental and clinical vaccine data analysis and vaccine literature mining. Overall, VO standardizes vaccine modeling and representation and greatly supports vaccine AI research in the Semantic Web era.

bioinformatics↗

VaxKG: Integrating The Vaccine Ontology And VIOLIN For Advanced Vaccine Queries And LLM-Powered Chat Systems

Vaccine research faces challenges in integrating diverse biomedical datasets. While the Vaccine Investigation and Online Information Network (VIOLIN) provides comprehensive vaccine data, implemented in traditional relational models limit complex analysis. Similarly, the Vaccine Ontology (VO) offers standardized semantic frameworks but lacks comprehensive empirical data. This study addresses these limitations by developing the Vaccine Knowledge Graph (VaxKG) that integrates VIOLINs dataset with VOs standardized terminology. Using Neo4j, we transformed 12 core VIOLIN tables into a graph structure enriched with VO concepts. The resulting knowledge graph comprises 28,123 VIOLIN data nodes and 101,282 VO resource nodes, connected by 412,865 relationships. Our comparative analysis of Brucella and Influenza vaccines demonstrates VaxKGs ability to enable complex semantic queries and reveal insights unavailable from either resource alone. We further demonstrate VaxKGs utility through VaxChat, a large language model application that leverages the VaxKG as Retrieval-Augmented Generation (RAG) for intuitive vaccine information access.

bioinformatics↗