Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.06.10.658658

Multi-agent AI System for High Quality Metadata Curation at Scale

Abstract

High-quality metadata is essential for downstream AI applications, yet metadata curation remains a persistent bottleneck in biomedical research. The core challenge lies in balancing quality and scalability. Manual curation delivers high quality, reliable annotations but is time-intensive and non-scalable; automated approaches, including those based on natural language processing, can scale but often fall short in accuracy and completeness. This tradeoff is particularly severe in public datasets, where metadata is frequently distributed across multiple sources for a single dataset. A notable example includes multi-omics datasets available on GEO and similar public repositories, which are often accompanied by associated publications. These supplementary sources can be leveraged to substantially enhance metadata quality, thereby supporting downstream applications such as classifier model development or the identification of biologically relevant cohorts, etc. We present a first-in-class multi-agent AI system that bridges the gap, achieving both high quality and scalability in metadata curation. Built on large-language models (LLMs), our system orchestrates a set of specialized agents that collaboratively extract, normalize, and infer critical metadata fields such as tissue, disease, cell line, sampling site, demographics, and experimental context from GEO entries, associated publications. A central orchestrator agent delegates tasks such as data retrieval, document parsing, ontology mapping, and context inference to expert sub-agents, enabling scalable, end-to-end automation. Applied to GEO data sets, a notoriously difficult metadata domain, our system achieves a 93% recall on average across 23 key fields (all original terms, normalized terms, and corresponding ontology ids) that include information about disease, tissue, treatment, donor-related information, outperforming existing automated baselines and approaches expert level quality. We also present a system that can easily scale to curate hundreds of metadata fields of interest with similar precision. This work demonstrates that an LLM-based multi-agent architecture can overcome traditional trade-offs in metadata curation, enabling both precision and scale, and offers a promising path forward for curating large public biomedical repositories for downstream AI applications.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mondal, R., Dhruw, N. K., Sen, M., Sengupta, S., Maity, W., Palapetta, S., Jha, A.. 2025-06-11. Multi-agent AI System for High Quality Metadata Curation at Scale. https://doi.org/10.1101/2025.06.10.658658

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

Are all vet schools equal? An exploration of postgraduate qualification attainment by alumni cohorts of seven UK veterinary schools.

Abstract & IntroductionDo graduates from the different vet schools attain clinical and academic postgraduate qualifications (certificates, diplomas, masters, PhD and Fellowship) at the same rates? This is important because leadership and progress within the veterinary profession, as in human medicine, comes from advancing the frontier of our knowledge through research and clinical specialisation. Postgraduate qualifications are essential training for both and are therefore useful metrics to measure. In this study the Royal College of Veterinary Surgeons annual registers of qualified vets from 2000-2021 were analysed. Significant and substantial differences in the proportion of graduates from different universities attaining postgraduate qualifications were observed. Whilst associations identified by this analysis cannot prove causation, they do strongly suggest the wide range of university-specific factors, such as student selection criteria, teaching methods, curriculum design and assessments which contribute to the culture and ethos of the institution have an impact on the career trajectory of their graduates.

scientific communication and education↗

Fine-Grained Detection of AI-Generated Writing in the Biomedical Literature

Generative AI systems are rapidly being integrated into scientific workflows, yet the specific ways in which AI-generated prose appears in published literature remain poorly characterized. Here, we use Pangram, a transformer-based detector optimized for adversarial paraphrasing, to analyze full-length biomedical research articles from 13 major journals. Papers published in 2021-2024 showed almost no detectable AI-generated text, whereas manuscripts published in 2025 exhibited a sharp increase, with 12.4% containing at least one localized passage classified as AI-written. AI usage was highly nonuniform across authors and geography: 32% of papers originating from South Korean institutions and 26% papers from Chinese institutions contained AI-generated passages, compared to 7.4% from U.S. institutions. In a focused case analysis, six labs that published fully AI-generated manuscripts also produced additional papers with extensive AI-generated segments. Journals likewise differed, with high-selectivity venues rarely containing AI-authored prose, while high-volume journals accounted for most AI-positive manuscripts. Together, these findings provide the first detailed empirical map of how and where AI-generated writing is entering the scientific literature, underscoring the need for clear norms and policies governing the use of generative AI in scientific communication.

scientific communication and education↗