bioRxiv · 10.64898/2026.09.17.752433
Leakage-controlled benchmarking reveals generalization limits of deep learning for protein-ligand binding affinity prediction
Abstract
To address widespread data leakage and inconsistent evaluation in protein-ligand affinity prediction, we introduce PLABench, a leakage-controlled and target-centric benchmarking framework that enables standardized comparison across sequence- and structure-based models. We benchmark nine deep learning methods across blind CASP16 targets, leakage-controlled ChEMBL35 sets, and established Davis and KIBA datasets under rigorous data split settings, standardizing structural input via AlphaFold3 to ensure fair comparison. Although pretrained structure-based models achieve the highest overall accuracy, they show severe target-dependent variability, and increasing structural fidelity from predicted to experimental conformations yields no consistent gains. Meanwhile, sequence-based approaches surpass some structure-based methods on select targets, and protein family-level evaluations reveal uneven performance across families and substantial inter-model complementarity obscured by aggregate metrics. These findings demonstrate that training scale and structural input alone cannot guarantee cross-target generalization, highlighting the need for context-aware interaction modeling. PLABench provides an extensible open-source platform to facilitate these developments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wang, L., Cheng, J.. 2026-09-23. Leakage-controlled benchmarking reveals generalization limits of deep learning for protein-ligand binding affinity prediction. https://doi.org/10.64898/2026.09.17.752433
Cite the original work for its findings. Save a collection to share your selection of sources.