bioRxiv · 10.64898/2026.05.05.723092
PromptBio-Bench: Benchmarking LLM-based Bioinformatics Agents for End-to-End Data Analysis
Abstract
Large language model (LLM)-based agents hold transformative potential for automating bioinformatics workflows; however, systematic evaluations of their capabilities remain limited, hindering a clear assessment of their readiness for real-world application. We introduce PromptBio-Bench, a comprehensive evaluation suite of 244 expert-curated tasks spanning bioinformatics and data science at varied difficulty levels, and an evaluation framework for structured file comparison and scoring against expert reference answer files. Evaluation of three state-of-the-art bioinformatics agents revealed comparable performance between Biomni and ToolsGenie, with all agents showing a marked decline in accuracy as task difficulty increased. As foundation models and agent frameworks continue to evolve, PromptBio-Bench provides a valuable benchmark infrastructure for systematically tracking progress in agentic bioinformatics.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guo, W., Zhang, M., Han, B., Ma, Y., Leng, Y., Hebbar, S., Zhou, X., Gu, W., Yang, X., Dhar, S.. 2026-05-08. PromptBio-Bench: Benchmarking LLM-based Bioinformatics Agents for End-to-End Data Analysis. https://doi.org/10.64898/2026.05.05.723092
Cite the original work for its findings. Save a collection to share your selection of sources.