bioRxiv · 10.1101/2024.10.23.619867
Utilizing data imbalance to enhance compound-protein interaction prediction models
Abstract
Identifying potential compounds for target proteins is crucial in drug discovery. Current compound-protein interaction prediction models concentrate on utilizing more complex features to enhance capabilities, but this often incurs substantial computational burdens. Indeed, this issue arises from the limited understanding of data imbalance between proteins and compounds, leading to insufficient optimization of protein encoders. Therefore, we introduce a sequence-based predictor named FilmCPI, designed to utilize data imbalance to learn proteins with their numerous corresponding compounds. FilmCPI consistently outperforms baseline models across diverse datasets and split strategies, and its generalization to unseen proteins becomes more pronounced as the datasets expand. Notably, FilmCPI can be transferred to unseen protein families with sequence-based data from other families, exhibiting its practicability. The effectiveness of FilmCPI is attributed to different optimization speeds for diverse encoders, elucidating optimization imbalance in compound-protein prediction models. Additionally, these advantages of FilmCPI do not depend on increasing parameters, aiming to lighten model design with data imbalance.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lin, W., FUNG, C. C. A.. 2024-10-25. Utilizing data imbalance to enhance compound-protein interaction prediction models. https://doi.org/10.1101/2024.10.23.619867
Cite the original work for its findings. Save a collection to share your selection of sources.