bioRxiv · 10.64898/2026.02.05.704016
Widespread data leakage inflates performance estimates in cancer drug response prediction
Abstract
Drug response prediction models guide biomarker discovery, yet their reported accuracy depends on rigorous cross-validation (CV). We show that supervised feature screening applied to all samples before CV, a widespread practice, introduces data leakage that systematically underestimates prediction error. Across 265 drugs and 1,462 cancer cell lines, leakage-free CV raises mean squared error by 16.6% on average, with near-zero feature overlap between leaked and corrected pipelines (mean Jaccard 0.18; 36.2% of drugs share no features). Despite selecting five times more features, the leaked pipeline recovers known drug targets at nearly the same rate as the corrected pipeline, indicating that the inflated feature sets capture statistical artifacts rather than biological signal. A code-level audit of 32 published methods (2017-2024) confirms leakage in 23 (72%), spanning five distinct modes and cited over 3,000 times. The magnitude of inflation from this single leakage mode is comparable to improvement margins typically claimed over elastic net baselines, raising the possibility that some reported advances reflect evaluation artifacts. We provide a leakage taxonomy, audit guide and reference implementation for leakage-free evaluation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Asiaee, A., Strauch, J., Azinfar, L., Pal, S., Pua, H. H., Long, J. P., Coombes, K. R.. 2026-02-08. Widespread data leakage inflates performance estimates in cancer drug response prediction. https://doi.org/10.64898/2026.02.05.704016
Cite the original work for its findings. Save a collection to share your selection of sources.