bioRxiv · 10.64898/2025.12.19.695582
Understanding the LLM-based gene embeddings
Abstract
Large language model (LLM)-derived gene embeddings, generated from brief NCBI gene descriptions, have shown strong performance in recent biological applications, yet the biological information they contain remains unclear. We evaluate these embeddings using a Gene Set Enrichment Analysis (GSEA)-based framework that treats each embedding dimension as a potential carrier of pathway-level information. OpenAIs embeddings recover over 93% of Hallmark and C2 pathways, with pathway signals distributed across many dimensions. Even embeddings generated from gene symbols alone recover more than 64% of pathways, indicating substantial prior biological knowledge embedded in the model. Comparing 11 small language models reveals that domain-specific models perform best with minimal input, but all models approach OpenAI-level coverage when given modest textual context. Collectively, these results show that LLM-derived embeddings encode unexpectedly extensive pathway-level information, supporting their use as lightweight, informative representations for downstream biological analysis.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cai, Y., Gan, D., Zhang, H., Li, J.. 2025-12-22. Understanding the LLM-based gene embeddings. https://doi.org/10.64898/2025.12.19.695582
Cite the original work for its findings. Save a collection to share your selection of sources.