Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.10.07.617112

Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating?

Abstract

There is a growing number of articles about conversational AI (i.e., ChatGPT) for generating scientific literature reviews and summaries. Yet, comparative evidence lags its wide adoption by many clinicians and researchers. We explored ChatGPTs utility for literature search from an end-user perspective through the lens of clinicians and biomedical researchers. We quantitatively compared basic versions of ChatGPTs utility against conventional search methods such as Google and PubMed. We further tested whether ChatGPT user-support tools (i.e., plugins, web-browsing function, prompt-engineering, and custom-GPTs) could improve its response across four common and practical literature search scenarios: (1) high-interest topics with an abundance of information, (2) niche topics with limited information, (3) scientific hypothesis generation, and (4) for newly emerging clinical practices questions. Our results demonstrated that basic ChatGPT functions had limitations in consistency, accuracy, and relevancy. User-support tools showed improvements, but the limitations persisted. Interestingly, each literature search scenario posed different challenges: an abundance of secondary information sources in high interest topics, and uncompelling literatures for new/niche topics. This study tested practical examples highlighting both the potential and the pitfalls of integrating conversational AI into literature search processes, and underscores the necessity for rigorous comparative assessments of AI tools in scientific research. Author SummaryAs generative Artificial Intelligence (AI) tools become increasingly functional, the promise of this technology is creating a wave of excitement and anticipation around the globe including the wider scientific and biomedical community. Despite this growing excitement, researchers seeking robust, reliable, reproducible, and peer-reviewed findings have raised concerns about AIs current limitations, particularly in spreading and promoting misinformation. This emphasizes the need for continued discussions on how to appropriately employ AI to streamline the current research practices. We, as members of the scientific community and also end-users of conversational AI tools, seek to explore practical incorporations of AI for streamlining research practices. Here, we probed text-based research tasks--scientific literature mining-- can be outsourced to ChatGPT and to what extent human adjudication might be necessary. We tested different models of ChatGPT as well as augmentations such as plugins and custom GPT under different contexts of biomedical literature searching. Our results show that though at present, ChatGPT does not meet the level of reliability needed for it to be widely adopted for scientific literature searching. However, as conversational AI tools rapidly advance (a trend highlighted by the development of augmentations in this article), we envision a time when ChatGPT can become a great time saver for literature searches and make scientific information easily accessible.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yip, R., Sun, Y. J., Bassuk, A., Mahajan, V. B.. 2024-10-08. Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating?. https://doi.org/10.1101/2024.10.07.617112

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Multi-Lab Testing of Early Preclinical Discoveries Identifies Promising Treatments

A fundamental challenge in drug development is the frequent failure of early laboratory research to translate into clinical benefit. One promising solution is to confirm findings from exploratory single-laboratory studies across multiple laboratories before clinical testing. We investigated this approach following the conduct of preclinical multi-laboratory studies across different fields of medicine. For this, we evaluated effect sizes, experimental rigor, and a set of criteria to identify determinants of confirmation success. When tested under increased rigor, only a fraction of multi-laboratory studies confirmed the initial results. The underlying effect size reduction was associated with outcome-relevant experimental differences between exploratory and confirmatory stages. In summary, multi-laboratory studies proved highly informative and served as an effective filter for promising treatments.

scientific communication and education↗

Technology-enhanced learning in undergraduate neuroscience education: tractography-based virtual dissection in psychology

Background: Neuroanatomy poses a significant challenge for Psychology students due to its spatial and conceptual complexity. Educational approaches that enhance the relevance and visualization of neuroanatomical content may improve students learning experiences. This study implemented a tractography-based activity focused on the virtual dissection of the arcuate fasciculus, a major white matter pathway, in undergraduate Psychology students and examined the relationships between perceived learning and students perceptions of utility, difficulty and handling, and organizational aspects of the activity. Methods: First-year undergraduate Psychology students participated in a two-session tractography-based activity combining instruction on white matter anatomy and diffusion tractography with a hands-on virtual dissection of the arcuate fasciculus using research-grade software routinely employed in neuroscience research. Following the activity, students completed an anonymous questionnaire assessing perceived learning, utility, difficulty and handling, and organizational aspects of the activity. Pearson correlations, multiple regression analyses, and relative importance analyses were performed. Results: Sixty-eight students completed the questionnaire. Students reported generally positive perceptions of the activity across the evaluated dimensions, with perceived learning receiving the highest mean score (M = 3.44, SD = .85). Perceived utility showed the strongest association with perceived learning (r = .64, p < .001). The regression model explained 41% of the variance in perceived learning (R2 = .41, adjusted R2 = .38, p < .001). Perceived utility was the only significant predictor in the model ({beta} = .59, p = .001), accounting for 67.9% of the explained variance. Conclusions: The findings support the feasibility of integrating authentic neuroimaging tools into undergraduate neuroanatomy teaching. Students who perceived the activity as more useful also reported higher perceived learning outcomes, with perceived utility emerging as the strongest predictor of perceived learning. In contrast, perceived difficulty and handling, and organizational aspects did not make significant independent contributions. These results suggest that students perceptions of educational relevance may play an important role in technology-enhanced STEM learning experiences.

scientific communication and education↗

A randomized trial of grant writing coaching groups: Qualitative interviews revealing key elements of intervention efficacy

Background Training in grant proposal writing is an essential component of professional development for academic scientists in the biomedical and behavioral sciences. Despite the expansion of inter- and intra-institutional grant writing coaching groups as an approach to honing these skills, specific features that enhance or limit coaching group effectiveness have not been rigorously studied. Methods Qualitative inerviews were conducted with a subset of early-career investigators (n=204 and coaches (n=36) engaged in a national, U.S. based, group-randomized trial of grant writing coaching groups to test the effects of two variables on submission and funding of national-level proposals: (1) coaching duration (regular/extended dose) and (2) mode of engaging a scientific advisor (someone with content-aligned expertise) in the coaching process. This report focuses on interviews conducted upon completion of the regular coaching dose - 5 months of biweekly, group-based coaching sessions to support active proposal writing. Interviews were designed to identify which coaching group elements were perceived to be the most critical. Transcribed interviews were analyzed using deductive (participants) or open (coaches) coding to identify themes. Results Triangulation of results from participant and coach interviews showed strong concordance that coaches, peers, and scientific advisors all played key roles in supporting intervention efficacy. Sufficient alignment of scientific fields and/or methodologies among group members was important, although breadth of perspectives was also valued. Other critical group features were skilled and well-organized coaches, detailed feedback, peer-to-peer support (technical and psychosocial), and clear expectations for group functioning. Factors attenuating impact included variation in participants' engagement and "readiness to write," within-group mismatches of expertise or grant mechanisms, and limited research support at some participants' home institutions. Conclusion This study identified key elements of successful grant writing coaching groups and potential barriers to their effectiveness, while yielding insights about tailoring this approach for individuals at different stages of proposal development.

scientific communication and education↗