Search bioRxiv⌕ Search

bioRxiv · 10.64898/2025.12.11.693290

Using GPT-4 to Automate the Generation of Lay Summaries for Cancer Publications

Abstract

BackgroundCancer research literature is often riddled with technical jargon that is not digestible to the average person. Individuals interested in research studies may want to contribute through patient partner engagement or sample donation but find the relevant literature overwhelming. Through the generation of lay summaries, previously inaccessible research papers become easier to comprehend, especially for patient partners or data donors. With large language models (LLMs) continuing to advance, so does their capability to summarize large texts. ObjectivesIn this study, we examined whether LLMs can produce lay summaries of scientific literature at-scale, while maintaining readability and accuracy to their source texts. MethodsWe developed a tool to generate lay summaries of open-access article abstracts and their full texts with GPT-4-Turbo. Prompt development aimed for a target 8th grade reading level assessed with Flesch-Kincaid Grade Level. Human-review metrics were used to evaluate readability and accuracy when generated using abstracts versus full text articles. ResultsThe average Flesch-Kincaid Grade Level Score was 7.13 for abstract-based summaries and 7.39 for full text-based summaries, indicating summaries at around 7th grade reading level. Human-review metrics showed these summaries were of similar readability and accuracy when generated using abstracts versus full text articles, with mean accuracy scores from human review of 7.09 vs 7.42 out of 10 respectively. Additionally, qualitative patient-based assessment indicated these summaries would encourage participation in research studies. ConclusionBy generating lay summaries for complex and lengthy research papers, their scientific information becomes accessible to a larger audience, including patient partners interested in contributing to cancer research. Summaries that are easy to understand will allow participants to make informed decisions about their involvement and appreciate the impact of their contributions if and when their results are published. Lay SummaryThis study explores if artificial intelligence (AI) can help make hard to read cancer research papers easier to understand for members of the public. ProblemWhen people donate cancer tissue samples or participate in research studies, they often want to know how their contributions are being used. However, scientific papers are full of technical language thats hard for most people to grasp. People in past studies have said this can make them less willing to take part in research. MethodsThe study created a computer program using AI (GPT-4-Turbo) to turn complex kidney cancer research papers into simple summaries. They tested whether the AI could summarize both short abstracts and full-length papers effectively. They aimed for summaries at a 6th to 8th-grade reading level. This was to follow Canadian and U.S. health communication guidelines. ResultsThe AI created 106 summaries. Computer measures showed the summaries were close to a 7th-grade reading level. Though, researchers had to tell the AI to write for a 2nd-grade audience to achieve this. Of note, summaries from short abstracts were just as accurate and readable as those from full papers. Eighteen volunteers, including five patient partners, reviewed the summaries and rated them for clarity and accuracy. They were rated at around 7 out of 10 points. All patient partners said these summaries would help them decide whether to join research studies and feel more informed about how their contributions matter. Why It MattersThis tool could help patients and donors better understand research without needing a science degree. When people can see how studies work, they are more likely to participate in future research. While patient partners emphasized the need for summaries to be accurate and reliable, this approach shows promise as a unique strategy to better connect the public with research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Purdie, E., Yu, T. J., Weile, J., Lemaire, D., Courtot, M.. 2025-12-12. Using GPT-4 to Automate the Generation of Lay Summaries for Cancer Publications. https://doi.org/10.64898/2025.12.11.693290

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Internal Grant Review: A Pre-Submission Program for Early-Career Clinical and Translational Researchers

The iTHRIV Scholars Mentored Career Development Program initiated an Internal Grant Review (IGR) program in 2020 for current and recently graduated Scholars seeking funding through K and R awards from the NIH. The IGR program is designed to replicate the NIH review process and provide Scholars the opportunity to receive valuable feedback on their applications prior to NIH submission. A key characteristic of the program is the integration of REDCap, enabling automation and year-round offering while improving tracking and reporting efforts and capabilities. Results of the program have been overall positive, both in proposal development and participant feedback. IGR is a sustainable, valuable resource for early-career faculty competing for limited resources in the pursuit to become independently funded clinical translational scientists.

scientific communication and education↗

Multi-Lab Testing of Early Preclinical Discoveries Identifies Promising Treatments

A fundamental challenge in drug development is the frequent failure of early laboratory research to translate into clinical benefit. One promising solution is to confirm findings from exploratory single-laboratory studies across multiple laboratories before clinical testing. We investigated this approach following the conduct of preclinical multi-laboratory studies across different fields of medicine. For this, we evaluated effect sizes, experimental rigor, and a set of criteria to identify determinants of confirmation success. When tested under increased rigor, only a fraction of multi-laboratory studies confirmed the initial results. The underlying effect size reduction was associated with outcome-relevant experimental differences between exploratory and confirmatory stages. In summary, multi-laboratory studies proved highly informative and served as an effective filter for promising treatments.

scientific communication and education↗

Technology-enhanced learning in undergraduate neuroscience education: tractography-based virtual dissection in psychology

Background: Neuroanatomy poses a significant challenge for Psychology students due to its spatial and conceptual complexity. Educational approaches that enhance the relevance and visualization of neuroanatomical content may improve students learning experiences. This study implemented a tractography-based activity focused on the virtual dissection of the arcuate fasciculus, a major white matter pathway, in undergraduate Psychology students and examined the relationships between perceived learning and students perceptions of utility, difficulty and handling, and organizational aspects of the activity. Methods: First-year undergraduate Psychology students participated in a two-session tractography-based activity combining instruction on white matter anatomy and diffusion tractography with a hands-on virtual dissection of the arcuate fasciculus using research-grade software routinely employed in neuroscience research. Following the activity, students completed an anonymous questionnaire assessing perceived learning, utility, difficulty and handling, and organizational aspects of the activity. Pearson correlations, multiple regression analyses, and relative importance analyses were performed. Results: Sixty-eight students completed the questionnaire. Students reported generally positive perceptions of the activity across the evaluated dimensions, with perceived learning receiving the highest mean score (M = 3.44, SD = .85). Perceived utility showed the strongest association with perceived learning (r = .64, p < .001). The regression model explained 41% of the variance in perceived learning (R2 = .41, adjusted R2 = .38, p < .001). Perceived utility was the only significant predictor in the model ({beta} = .59, p = .001), accounting for 67.9% of the explained variance. Conclusions: The findings support the feasibility of integrating authentic neuroimaging tools into undergraduate neuroanatomy teaching. Students who perceived the activity as more useful also reported higher perceived learning outcomes, with perceived utility emerging as the strongest predictor of perceived learning. In contrast, perceived difficulty and handling, and organizational aspects did not make significant independent contributions. These results suggest that students perceptions of educational relevance may play an important role in technology-enhanced STEM learning experiences.

scientific communication and education↗