bioRxiv · 10.1101/2025.07.12.664517
Evaluating large language models in biomedical data science challenges through a classroom experiment
Abstract
Large language models have shown remarkable capabilities in algorithm design, but their effectiveness in solving data science challenges remains poorly understood. We conducted a classroom experiment in which graduate students used large language models (LLMs) to solve biomedical data science challenges on Kaggle. While their submissions did not top the leaderboards, their prediction scores were often close to those of leading human participants. LLMs frequently recommended gradient boosting methods, which were associated with better performance. Among prompting strategies, self-refinement, where the LLM improves its own initial solution, was the most effective, a result validated using additional LLMs. These findings demonstrate that LLMs can design competitive machine learning solutions, even when used by non-experts.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ma, H., BIOSTAT 824 Student Consortium,, Ji, Z.. 2025-07-17. Evaluating large language models in biomedical data science challenges through a classroom experiment. https://doi.org/10.1101/2025.07.12.664517
Cite the original work for its findings. Save a collection to share your selection of sources.