Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.01.23.701317

Set-up, validation, evaluation, and cost-benefit analysis of an AI-assisted assessment of responsible research practices in a sample of life science publications

Abstract

The (semi-)automated screening of publications for diverse quality and transparency criteria is at the core of systematic literature assessment. Typically, the assessment process involves two initial reviewers and one additional reviewer for cases that require reconciliation. Here, we explore to what extent this process can be assisted by Large Language Models (LLMs). Specifically, whether LLMs are capable of assessing responsible research practices (RRPs) in scientific papers in a robust way. We employed proprietary LLMs to assess an initial set of 37 papers across ten RRPs. The same papers were also reviewed by three human reviewers. We iteratively redesigned prompts to increase model accuracy compared to human ratings which we treated as the gold standard. The resulting pipeline was validated on an additional set of 15 papers. We show that LLM accuracy is comparable to single human reviewer performance (90% for LLM vs 86% for a single human reviewer). However, performance strongly depended on the specific RRPs with accuracy ranging from 40% to 100%. LLMs exhibited an affirmative bias, making more errors when practices were not reported in the papers. Overall, we show how such an approach potentially replaces one human reviewer, enabling AI-assisted assessment of research papers. We discuss how dataset imbalances, validation procedures, and implementation time limit the broad applicability of such approaches. Through this, we develop initial guidance on the utility of proprietary LLMs in evidence synthesis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kniffert, S., Kathoefer, B., Emprechtinger, R., Pellegrini, P., Funk, E. M., Dhamrait, I. S., Zang, Y., Bornmueller, A., Toelch, U.. 2026-02-02. Set-up, validation, evaluation, and cost-benefit analysis of an AI-assisted assessment of responsible research practices in a sample of life science publications. https://doi.org/10.64898/2026.01.23.701317

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Internal Grant Review: A Pre-Submission Program for Early-Career Clinical and Translational Researchers

The iTHRIV Scholars Mentored Career Development Program initiated an Internal Grant Review (IGR) program in 2020 for current and recently graduated Scholars seeking funding through K and R awards from the NIH. The IGR program is designed to replicate the NIH review process and provide Scholars the opportunity to receive valuable feedback on their applications prior to NIH submission. A key characteristic of the program is the integration of REDCap, enabling automation and year-round offering while improving tracking and reporting efforts and capabilities. Results of the program have been overall positive, both in proposal development and participant feedback. IGR is a sustainable, valuable resource for early-career faculty competing for limited resources in the pursuit to become independently funded clinical translational scientists.

scientific communication and education↗

Multi-Lab Testing of Early Preclinical Discoveries Identifies Promising Treatments

A fundamental challenge in drug development is the frequent failure of early laboratory research to translate into clinical benefit. One promising solution is to confirm findings from exploratory single-laboratory studies across multiple laboratories before clinical testing. We investigated this approach following the conduct of preclinical multi-laboratory studies across different fields of medicine. For this, we evaluated effect sizes, experimental rigor, and a set of criteria to identify determinants of confirmation success. When tested under increased rigor, only a fraction of multi-laboratory studies confirmed the initial results. The underlying effect size reduction was associated with outcome-relevant experimental differences between exploratory and confirmatory stages. In summary, multi-laboratory studies proved highly informative and served as an effective filter for promising treatments.

scientific communication and education↗

Technology-enhanced learning in undergraduate neuroscience education: tractography-based virtual dissection in psychology

Background: Neuroanatomy poses a significant challenge for Psychology students due to its spatial and conceptual complexity. Educational approaches that enhance the relevance and visualization of neuroanatomical content may improve students learning experiences. This study implemented a tractography-based activity focused on the virtual dissection of the arcuate fasciculus, a major white matter pathway, in undergraduate Psychology students and examined the relationships between perceived learning and students perceptions of utility, difficulty and handling, and organizational aspects of the activity. Methods: First-year undergraduate Psychology students participated in a two-session tractography-based activity combining instruction on white matter anatomy and diffusion tractography with a hands-on virtual dissection of the arcuate fasciculus using research-grade software routinely employed in neuroscience research. Following the activity, students completed an anonymous questionnaire assessing perceived learning, utility, difficulty and handling, and organizational aspects of the activity. Pearson correlations, multiple regression analyses, and relative importance analyses were performed. Results: Sixty-eight students completed the questionnaire. Students reported generally positive perceptions of the activity across the evaluated dimensions, with perceived learning receiving the highest mean score (M = 3.44, SD = .85). Perceived utility showed the strongest association with perceived learning (r = .64, p < .001). The regression model explained 41% of the variance in perceived learning (R2 = .41, adjusted R2 = .38, p < .001). Perceived utility was the only significant predictor in the model ({beta} = .59, p = .001), accounting for 67.9% of the explained variance. Conclusions: The findings support the feasibility of integrating authentic neuroimaging tools into undergraduate neuroanatomy teaching. Students who perceived the activity as more useful also reported higher perceived learning outcomes, with perceived utility emerging as the strongest predictor of perceived learning. In contrast, perceived difficulty and handling, and organizational aspects did not make significant independent contributions. These results suggest that students perceptions of educational relevance may play an important role in technology-enhanced STEM learning experiences.

scientific communication and education↗