bioRxiv · 10.1101/2024.05.14.594226
ProteinCLIP: enhancing protein language models with natural language
Abstract
Language models have enabled a new era of biological sequence modeling. However, extracting meaningful sequence-level embeddings from these models remains challenging. In this work, we introduce ProteinCLIP, which applies contrastive learning between a proteins amino acid sequence and curated text describing its function. ProteinCLIP thus learns to take a pre-trained protein language models sequence embedding and refines it produce a function-centric embedding. We show that this embedding space yields sequence representations that enable state-of-the-art performance across a variety of important yet challenging tasks in the study of proteins - from predicting protein protein interactions to accurately detecting homologous proteins despite low sequence similarity. More broadly, ProteinCLIP demonstrates the effectiveness of multi-modal learning in biological contexts, and how such strategies can help isolate key signals from large models and further improve their utility.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wu, K. E., Chang, H., Zou, J.. 2024-05-17. ProteinCLIP: enhancing protein language models with natural language. https://doi.org/10.1101/2024.05.14.594226
Cite the original work for its findings. Save a collection to share your selection of sources.