bioRxiv · 10.64898/2026.01.26.701583
Preventing Data Leakage in Neural Decoding
Abstract
Neural decoding is a widely-used machine learning technique for investigating how behavior, perception and cognition are represented in neural activity. However without careful application data leakage can occur, where information from the test set contaminates the training set, leading to biased estimates of decoding performance and potentially invalidating biological conclusions. Here we show that leakage can be introduced in neural decoding studies by common supervised and unsupervised preprocessing steps, including dimensionality reduction. We reveal that in some cases leakage can paradoxically decrease decoding performance relative to unbiased estimates, and we provide theoretical analyses explaining how this occurs. We also show that, for autocorrelated neural time series, randomized k-fold cross-validation can dramatically overstate decoding performance. Based on these findings, we provide detailed recommendations for avoiding data leakage in neural decoding.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wong, R., Zhu, S. I., McCullough, M. H., Goodhill, G. J.. 2026-01-27. Preventing Data Leakage in Neural Decoding. https://doi.org/10.64898/2026.01.26.701583
Cite the original work for its findings. Save a collection to share your selection of sources.