Dopamine reveals adaptive learning of actions representation
Flexible decision-making requires not only updating values, but redefining which features constitute an action in a given context. We recorded nucleus accumbens (NAc) dopamine release while mice navigated a three-target intracranial self-stimulation foraging task in which outcomes were evaluated under three distinct reward delivery rules. Despite a constant motor repertoire, dopamine transients reorganized across contingencies and generalized linear models revealed context-dependent dopamine signal reflecting action direction, recent outcome-history, or target identity. Reinforcement-learning model comparison showed that these signatures are best explained by distinct reward prediction errors (RPEs) defined over different state-action representations, rather than a single fixed model-free scheme. A single deep reinforcement-learning agent trained by temporal-difference learning, recapitulated both the rule-specific policies and the corresponding dopamine signature. These results identify NAc dopamine as a dynamic readout of representation learning, remapping prediction errors onto the task features that define successful action as contingencies change.