bioRxiv · 10.64898/2026.08.18.745260
PerturbTrace: Evaluating Feedback Use by AI Co-Scientist Agents in Perturbation Discovery
Abstract
Recent advances in AI co-scientists have brought LLM agents into closed-loop experimental design. However, whether these agents use feedback from earlier rounds to revise subsequent experimental decisions remains unclear. We address this question with PerturbTrace, which evaluates each round-to-round transition through Feedback-to-State, State-to-Action, and Action-to-Outcome. These stages assess whether feedback is reflected in the agent's rationale and perturbation-selection strategy, whether the stated strategy guides the next perturbation batch, and whether that batch yields more hits than expected under random sampling. We evaluate four LLM agents on 17 screen-derived tasks and compare them with random selection, active learning, and LLM-guided Bayesian optimization baselines. Each agent outperforms the strongest non-agent method on at least 15 of the 17 tasks, yet controlled evaluations across six tasks show no consistent advantage from true feedback over random or no feedback. Among 576 transitions under true or random feedback, only 43 (7.5%) complete the full Feedback-State-Action-Outcome sequence, including 25 under random feedback. These findings show that high final recall does not necessarily indicate effective feedback use. They also highlight the need to evaluate closed-loop scientific agents by both their discovery performance and whether feedback changes their subsequent decisions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yu, C., Liu, S., Qiao, G., Luo, M., Xiang, Y., Xu, Z.. 2026-08-20. PerturbTrace: Evaluating Feedback Use by AI Co-Scientist Agents in Perturbation Discovery. https://doi.org/10.64898/2026.08.18.745260
Cite the original work for its findings. Save a collection to share your selection of sources.