bioRxiv · 10.64898/2026.02.17.706287
Reliable Evaluation and Learning in Multi-input Biological Association Prediction
Abstract
Multi-input association prediction is central to many key problems in computational biology, spanning tasks from drug-target and protein-protein interactions to higher-order challenges such as drug synergy modeling and MHC-peptide-TCR binding. Yet widely used benchmarks often overestimate performance by enabling models to exploit degree ratio shortcuts, while alternative out-of-distribution splits are overly restrictive and impractical. Here we introduce an entity-balanced evaluation framework that systematically neutralizes shortcut signals by balancing positive and negative associations at the entity level. This enables fairer assessments that reflect genuine relational learning and extend naturally from pairwise to multi-entity problems. We further present UnbiasNet, a model-agnostic training strategy that cycles through diverse entity-balanced sub-training sets, removing access to degree ratio bias and enhancing robustness. Applied to drug-target interaction and drug synergy prediction, our framework reveals the extent of shortcut reliance in existing methods while enabling consistent identification of meaningful biological associations, thereby setting a rigorous foundation for future methodological progress.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ahmadian Moghadam, S., Montazeri, H.. 2026-02-18. Reliable Evaluation and Learning in Multi-input Biological Association Prediction. https://doi.org/10.64898/2026.02.17.706287
Cite the original work for its findings. Save a collection to share your selection of sources.