A novel critic signal in identified midbrain dopaminergic neurons of mice training inoperant tasks
Classically, midbrain dopaminergic neuron activity is triggered by unexpected rewards, then, upon learning, by reward-predictive conditioned stimuli. When expected rewards are withheld, firing is inhibited. This activity occurs too late to directly affect the neuronal circuitry underlying decision-making, inspiring the development of temporal difference (TD) reinforcement learning models. To test for more timely critical feedback during decision-making and learning, we recorded optogenetically identified dopaminergic, putative GABAergic and other neurons of the ventral tegmental area (VTA) and substantia nigra pars compacta in mice training in visual and olfactory discrimination tasks. The mice often adhered to unrewarded and untrained task strategies (e.g., spatial alternation) rather than making random choices. In order to probe for reward/punishment predictive activity, a delay was imposed between nose-poke choices and signals for reward or punishment. As animals performed below criterion levels, dopaminergic and other neurons firing rates signaled correct versus incorrect choices immediately after choices, but prior to the onset of trial outcome signals. Thus, this activity signaled the rewarded rule even as the mice performed other unrewarded strategies. Putative GABAergic neurons fired during nose-poke choices, potentially reducing network activity prior to reward prediction signals. This reward predictive activity could serve as a critic signal expressed immediately after choices are made, priming the network for canonical DA reward/punishment activity, facilitating network functional modifications. This is consistent with a role for dopamine in arbitration between brain modules to choose among diverse strategies during goal-directed behavior. These findings suggest extensions of theoretical formulations interpreting dopaminergic neuronal activity. Significance statementModels of dopaminergic influence on circuit modifications during learning evoke mechanisms dealing with the delay between neural activity leading up to choices, and when the reinforcing outcome actually occurs. Here, as mice trained in sensory discrimination tasks with a delay between behavioral responses and reinforcement signals, they performed several unrewarded behavioral strategies. Simultaneously, dopaminergic nuclei neurons instead reflected the current task rule, predicting whether the choice was correct or not, providing an immediate "critic" signal prior to canonical trial outcome signals. This provides evidence for brain mechanisms to overcome innate or acquired habits to perform behaviors optimizing positive outcomes.