Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. II. Antibody Properties, Lipid-RNA Interactions
Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their utility for real-world industrial research remains insufficiently characterized. Extending the analysis presented in the first paper of this series, we evaluated the same five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and the AI Scientist-v2 from Sakana AI) on two more projects of high practical importance for biopharmaceutical development: predicting antibody developability properties with the use of pretrained protein language model embeddings, and modeling non-covalent lipid-RNA interactions in lipid nanoparticles with all-atom molecular dynamics (MD) simulations. The AI frameworks again showed genuine strengths, including unprompted identification of subtle methodological issues, successful use of pretrained protein embeddings, and consistent reporting of p-values and confidence intervals often absent from the original papers. However, no framework approached the scope of the original studies, and severe failures and hallucinations were observed. Our results confirm and extend the conclusion of the first paper that real published research from pharmaceutical companies that we tried to reproduce proved to be considerably harder for current AI frameworks than standard benchmarks suggest.