bioRxiv · 10.1101/2024.06.24.600176
GeneRAG: Enhancing Large Language Models with Gene-Related Task by Retrieval-Augmented Generation
Abstract
Large Language Models (LLMs) like GPT-4 have revolutionized natural language processing and are used in gene analysis, but their gene knowledge is incomplete. Fine-tuning LLMs with external data is costly and resource-intensive. Retrieval-Augmented Generation (RAG) integrates relevant external information dynamically. We introduce GO_SCPLOWENEC_SCPLOWRAG, a frame-work that enhances LLMs gene-related capabilities using RAG and the Maximal Marginal Relevance (MMR) algorithm. Evaluations with datasets from the National Center for Biotechnology Information (NCBI) show that GO_SCPLOWENEC_SCPLOWRAG outperforms GPT-3.5 and GPT-4, with a 39% improvement in answering gene questions, a 43% performance increase in cell type annotation, and a 0.25 decrease in error rates for gene interaction prediction. These results highlight GO_SCPLOWENEC_SCPLOWRAGs potential to bridge a critical gap in LLM capabilities for more effective applications in genetics.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lin, X., Deng, G., Li, Y., Ge, J., Ho, J. W. K., Liu, Y.. 2024-06-28. GeneRAG: Enhancing Large Language Models with Gene-Related Task by Retrieval-Augmented Generation. https://doi.org/10.1101/2024.06.24.600176
Cite the original work for its findings. Save a collection to share your selection of sources.