bioRxiv · 10.1101/2024.05.10.592927
LucaOne: Generalized Biological Foundation Model with Unified Nucleic Acid and Protein Language
Abstract
AbstractThe language of biology, encoded in DNA, RNA, and proteins, forms the foundation of life but remains challenging to decode due to its complexity. Traditional computational methods often struggle to integrate information across these molecules, limiting a comprehensive understanding of biological systems. Advances in Natural Language Processing (NLP) with pre-trained models offer new possibilities for interpreting biological language. Here, we introduce LucaOne, a pre-trained foundation model trained on nucleic acid and protein sequences from 169,861 species. Through large-scale data integration and semisupervised learning, LucaOne demonstrates an understanding of key biological principles, such as DNA-Protein translation. Using few-shot learning, it effectively comprehends the central dogma of molecular biology and performs competitively on tasks involving DNA, RNA, or protein inputs. Our results highlight the potential of unified foundation models to address complex biological questions, providing an adaptable framework for bioinformatics research and enhancing the interpretation of lifes complexity.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
He, Y., Fang, P., Shan, Y., Pan, Y., Wei, Y., Chen, Y., Liu, Y., Zeng, Z., Zhou, Z., Zhu, F., Holmes, E. C., Ye, J., Li, J., Shu, Y., Shi, M., Li, Z.. 2024-05-14. LucaOne: Generalized Biological Foundation Model with Unified Nucleic Acid and Protein Language. https://doi.org/10.1101/2024.05.10.592927
Cite the original work for its findings. Save a collection to share your selection of sources.