bioRxiv · 10.1101/2024.08.13.607829
CLR_ESP: Improved enzyme-substrate pair prediction using contrastive learning
Abstract
To reduce the cost of experimental characterization of the potential substrates for enzymes, machine learning prediction model offers an alternative solution. Pretrained language models, as powerful approaches for protein and molecule representation, have been employed in the development of enzyme-substrate prediction models, achieving promising performance. In addition to continuing improvements in language models, effectively fusing encoders to handle multimodal prediction tasks is critical for further enhancing model performance using available representation methods. Here, we present FusionESP, a multimodal architecture that integrates protein and chemistry language models with a newly designed contrastive learning strategy for predicting enzyme-substrate pairs. Our best model achieved state-of-the-art performance with an accuracy of 94.77% on independent test data and exhibited better generalization capacity while requiring fewer computational resources and training data, compared to previous studies of finetuned encoder or employing more encoders. It also confirmed our hypothesis that embeddings of positive pairs are closer to each other in high-dimension space, while negative pairs exhibit the opposite trend. The proposed architecture is expected to be further applied to enhance performance in additional multimodality prediction tasks in biology. A user-friendly web server of FusionESP is established and freely accessible at https://rqkjkgpsyu.us-east-1.awsapprunner.com/.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Du, Z., Fu, W., Guo, X., Caragea, D., Li, Y.. 2024-08-16. CLR_ESP: Improved enzyme-substrate pair prediction using contrastive learning. https://doi.org/10.1101/2024.08.13.607829
Cite the original work for its findings. Save a collection to share your selection of sources.