bioRxiv · 10.1101/2025.10.10.681616
Optimal composition of multiple value functions for dopamine-mediated efficient, safe and stable learning
Abstract
AO_SCPLOWBSTRACTC_SCPLOWThe seminal reward prediction error account of dopamine has been highly successful, but faces several key challenges. Most notable are the difficulty of learning multiple rewards simultaneously, inefficient on-policy learning, and accounting for the heterogeneous striatal responses observed across and within striatal targets. Here we address these issues with a normative entropy-regularised reinforcement-learning framework. We propose that dopamine optimises not just cumulative rewards, but a reward value function augmented by a penalty for deviating from a default behavioural policy. In simulations, this off-policy formulation provides a principled solution to composing multiple reward values, avoids the interference and unintended unlearning seen in multi-objective on-policy methods when priorities change, and adapts more efficiently than standard alternatives in environments with non-stationary rewards. More broadly, the framework offers a unified account of dopamine heterogeneity between and within striatal targets, including a normative way to understand why aversive and action prediction errors may coexist in the tail of the striatum. Together, these results suggest that dopamine-mediated learning may be better captured by prediction errors in composable, entropy-regularised value functions than by a single broadcast prediction error, and offer testable predictions for future experiments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mahajan, P., Seymour, B.. 2025-10-10. Optimal composition of multiple value functions for dopamine-mediated efficient, safe and stable learning. https://doi.org/10.1101/2025.10.10.681616
Cite the original work for its findings. Save a collection to share your selection of sources.