Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.06.01.657285

Active learnings impact on student course performance in STEM varies by type and intensity

Abstract

We updated a recent meta-analysis of active learnings impact on student achievement in undergraduate STEM courses by following the same protocol to evaluate studies published from 2010-2017. We screened 1659 papers, coded 1294, and found 210 that met five pre-established inclusion criteria and six pre-established criteria for methodological quality. After further dropping 76 studies with no exam scores data, 134 of these studies contained data on student performance on identical or equivalent exams. We found that on average, active learnings effect size on exam scores was 0.519 {+/-} 0.049, meaning that when students are in active learning classes, they perform roughly half a standard deviation higher on an identical exam. Funnel plots and sensitivity analyses indicated that these results were not due to sampling bias. Active learning had a positive impact on student outcomes regardless of class size, course level, or STEM discipline, though there was heterogeneity in the effects. All of these results are very similar when compared to earlier meta-analyses, however increased resolution in the studies analyzed here revealed two novel results. First, student performance was significantly better in courses that employed high-intensity active learning, defined as students being on task at least two-thirds of class time, versus lower-intensities. Additionally, there was significant heterogeneity in efficacy across different types of active learning employed. These results suggest that most, if not all types of active learning are effective, and that when innovating in their classes, instructors should continually work to increase active learning intensity. We urge caution in interpreting the results on active learning types, however, and propose a preliminary framework for making more-sophisticated and reliable analyses of variation in course design. Finally, the evidence presented here for active learnings impact on student outcomes creates a strong foundation for faculty professional development and administration.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xu, S., Velasco, V., Hill, M. J., Tran, E., Agrawal, S., Arroyo, E. N., Behling, S., Chambwe, N., Cintron, D. L., Cooper, J. D., Dunster, G., Grummer, J. A., Hennessey, K., Hsiao, J., Iranon, N., Jones, L., Jordt, H., Keller, M., Lacey, M. E., Littlefield, C. E., Lowe, A., Newman, S., Okolo, V., Olroyd, S., Peecook, B. R., Pickett, S. B., Slager, D. L., Caviedes-Solis, I. W., Stanchak, K. E., Sundaravaradan, V., Valdebenito, C., Williams, C. R., Zinsli, K. A., Freeman, S., Theobald, E. J.. 2025-06-02. Active learnings impact on student course performance in STEM varies by type and intensity. https://doi.org/10.1101/2025.06.01.657285

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

A strong start for sustained success: inclusivity through a national group mentorship program for first-year graduate students

In the United States, STEM graduate programs and workforce do not represent the demographics of the population. Obstacles, including a lack of transparency, community, and accessible information in navigating academia, disproportionately affect students from underserved backgrounds. Peer mentoring networks can address these disparities. Here, we describe Cientifico Latino, Inc.s Graduate Student Engagement and Community (CL-GSEC) program, a nationwide, group-based peer mentorship program that has served first-year graduate students across the U.S., especially those from underserved backgrounds. Surveys indicate CL-GSEC positively impacts the first-year graduate experience. We highlight key program features, challenges, and insights, such as financial strains faced by first-year graduate students. We offer suggestions for how faculty and departments can better support students during this critical early stage of graduate training. We hope that reporting on CL-GSECs program structure, evaluations, and findings will guide educational leaders in expanding programming for junior graduate students.

scientific communication and education↗

Biodesign Buddy: Integrating Generative Artificial Intelligence in Academic Biodesign

Biodesign is an interdisciplinary research domain that incorporates principles from design and the life sciences to develop new systems, processes, and objects. Collegiate biodesign educators face unique pedagogical challenges, including an absence of relevant scholarship on curriculum design and instructional best practices for cultivating student scientific literacy. These difficulties may be overcome with newly available technologies, like generative AI systems, that enable personalized learning through domain-specific semantic spaces. This article examines the instructional value of one such domain-specific LLM, Biodesign Buddy, through a mixed-methods analysis of an eight-week study involving 64 students participating in an international biodesign competition. Results indicate strong support for integrating AI into biodesign coursework. Surveys captured attitudes toward AI, scientific literature, and learning experiences to assess AIs impact on learning outcomes. Findings suggest that integrating AI into biodesign pedagogy can meaningfully redress conceptual issues in biodesign while informing broader debates on AIs role in higher education. Impact StatementThis article introduces Biodesign Buddy, a domain-specific generative AI system for collegiate biodesign education, and reports on its exploratory deployment, offering design principles and preliminary findings to inform the development of AI-supported pedagogies for interdisciplinary biodesign instruction.

scientific communication and education↗