Search bioRxivSearch

Biology subjects

Botvinick, M. M.

Publications and source records attributed to Botvinick, M. M..

8 recordsLinked to original sources

Episodic Control as Meta-Reinforcement Learning

Recent research has placed episodic reinforcement learning (RL) alongside model-free and model-based RL on the list of processes centrally involved in human reward-based learning. In the present work, we extend the unified account of model-free and model-based RL developed by Wang et al. (2018) to further integrate episodic learning. In this account, a generic model-free \"meta-learner\" learns to deploy and coordinate among all of these learning algorithms. The meta-learner learns through brief encounters with many novel tasks, so that it learns to learn about new tasks. We show that when equipped with an episodic memory system inspired by theories of reinstatement and gating, the meta-learner learns to use the episodic and model-based learning algorithms observed in humans in a task designed to dissociate among the influences of various learning strategies. We discuss implications and predictions of the model.

neuroscience

Value Representations in Orbitofrontal Cortex Drive Learning, but not Choice

Humans and animals make predictions about the rewards they expect to receive in different situations. In formal models of behavior, these predictions are known as value representations, and they play two very different roles. Firstly, they drive choice: the expected values of available options are compared to one another, and the best option is selected. Secondly, they support learning: expected values are compared to rewards actually received, and future expectations are updated accordingly. Whether these different functions are mediated by different neural representations remains an open question. Here we employ a recently-developed multi-step task for rats that computationally separates learning from choosing. We investigate the role of value representations in the rodent orbitofrontal cortex, a key structure for value-based cognition. Electrophysiological recordings and optogenetic perturbations indicate that these representations do not directly drive choice. Instead, they signal expected reward information to a learning process elsewhere in the brain that updates choice mechanisms.

neuroscience

Subgoal- and Goal-Related Prediction Errors in Medial Prefrontal Cortex

A longstanding view of the organization of human and animal behavior holds that behavior is hierarchically organized, meaning that it can be understood as directed towards achieving superordinate goals through subordinate goals, or subgoals. For example, the superordinate goal of making coffee can be broken down as accomplishing a series of subgoals, namely boiling water, grinding coffee, pouring cream, etc. Learning and behavioral adaptation depend on prediction-error signals, which have been observed in ventral striatum (VS) and medial prefrontal cortex (mPFC). In past work, we have shown that prediction error signals (PEs) can be linked not only to superordinate goals, but also to subgoals. Here we present two functional magnetic resonance imagining experiments that replicate and extend these findings. In the first experiment, we replicated the finding that mPFC signals subgoal-related PEs, independently of goal PEs. Together with our past work, this experiment reveals that BOLD responses to PEs in mPFC are unsigned. In the second experiment, we showed that when a task involves both goal and subgoal PEs, mPFC shows only goal-related PEs, suggesting that context or attention can strongly impact hierarchical PE coding. Furthermore, we observed a dissociation between the coding of PEs in mPFC and VS. These experiments suggest that the mPFC selectively attends to information at different levels of hierarchy depending on the task context.

neuroscience

Dissociable neural mechanisms track evidence accumulation for selection of attention versus action

Decision-making is typically studied as a sequential process from the selection of what to attend (e.g., between possible tasks, stimuli, or stimulus attributes) to the selection of which actions to take based on the attended information. However, people often gather information across these levels in parallel. For instance, even as they choose their actions, they may continue to evaluate how much to attend other tasks or dimensions of information within a task. We scanned participants while they made such parallel evaluations, simultaneously weighing how much to attend two dynamic stimulus attributes and which response to give based on the attended information. Regions of prefrontal cortex tracked information about the stimulus attributes in dissociable ways, related to either the predicted reward (ventromedial prefrontal cortex) or the degree to which that attribute was being attended (dorsal anterior cingulate, dACC). Within dACC, adjacent regions tracked uncertainty at different levels of the decision, regarding what to attend versus how to respond. These findings bridge research on perceptual and value-based decision-making, demonstrating that people dynamically integrate information in parallel across different levels of decision making.\n\nNaturalistic decisions allow an individual to weigh their options within a particular task (e.g., how best to word the introduction to a paper) while also weighing how much to attend other tasks (e.g., responding to e-mails). These different types of decision-making have a hierarchical but reciprocal relationship: Decisions at higher levels inform the focus of attention at lower levels (e.g., whether to select between citations or email addresses) while, at the same time, information at lower levels (e.g., the salience of an incoming email) informs decisions regarding which task to attend. Critically, recent studies suggest that decisions across these levels may occur in parallel, continuously informed by information that is integrated from the environment and from ones internal milieu1,2.\n\nResearch on cognitive control and perceptual decision-making has examined how responses are selected when attentional targets are clearly defined (e.g., based on instruction to attend a stimulus dimension), including cases in which responding requires accumulating information regarding a noisy percept (e.g., evidence favoring a left or right response)3-7. Separate research on value-based decision-making has examined how individuals select which stimulus dimension(s) to attend in order to maximize their expected rewards8-11. However, it remains unclear how the accumulation of evidence to select high-level goals and/or attentional targets interacts with the simultaneous accumulation of evidence to select responses according to those goals (e.g., based on the perceptual properties of the stimuli). Recent work has highlighted the importance of such interactions to understanding task selection12-15, multi-attribute decision-making16-18, foraging behavior19-21, cognitive effort22,23, and self-control24-27.\n\nWhile these interactions remain poorly understood, previous research has identified candidate neural mechanisms associated with multi-attribute value-based decision-making11,28,29 and with selecting a response based on noisy information from an instructed attentional target3-5. These research areas have implicated the ventromedial prefrontal cortex (vmPFC) in tracking the value of potential targets of attention (e.g., stimulus attributes)8,11 and the dorsal anterior cingulate cortex (dACC) in tracking an individuals uncertainty regarding which response to select30-32. It has been further proposed that dACC may differentiate between uncertainty at each of these parallel levels of decision-making (e.g., at the level of task goals or strategies vs. specific motor actions), and that these may be separately encoded at different locations along the dACCs rostrocaudal axis32,33. However, neural activity within and across these prefrontal regions has not yet been examined in a setting in which information is weighed at both levels within and across trials.\n\nHere we use a value-based perceptual decision-making task to examine how people integrate different dynamic sources of information to decide (a) which perceptual attribute to attend and (b) how to respond based on the evidence for that attribute. Participants performed a task in which they regularly faced a conflict between attending the stimulus attribute that offered the greater reward or the attribute that was more perceptually salient (akin to persevering in writing ones paper when an enticing email awaits). We demonstrate that dACC and vmPFC track evidence for the two attributes in dissociable ways. Across these regions, vmPFC weighs attribute evidence by the reward it predicts and dACC weighs it by its attentional priority (i.e., the degree to which that attribute drives choice). Within dACC, adjacent regions differentiated between uncertainty at the two levels of the decision, regarding what to attend (rostral dACC) versus how to respond (caudal dACC).

neuroscience

The hippocampus as a predictive map

A cognitive map has long been the dominant metaphor for hippocampal function, embracing the idea that place cells encode a geometric representation of space. However, evidence for predictive coding, reward sensitivity, and policy dependence in place cells suggests that the representation is not purely spatial. We approach this puzzle from a reinforcement learning perspective: what kind of spatial representation is most useful for maximizing future reward? We show that the answer takes the form of a predictive representation. This representation captures many aspects of place cell responses that fall outside the traditional view of a cognitive map. Furthermore, we argue that entorhinal grid cells encode a low-dimensional basis set for the predictive representation, useful for suppressing noise in predictions and extracting multiscale structure for hierarchical planning.

neuroscience

Dorsal hippocampus plays a causal role in model-based planning

Planning can be defined as a process of action selection that leverages an internal model of the environment. Such models provide information about the likely outcomes that will follow each selected action, and their use is a key function underlying complex adaptive behavior. However, the neural mechanisms supporting this ability remain poorly understood. In the present work, we adapt for rodents recent advances from work on human planning, presenting for the first time a task for animals which produces many trials of planned behavior per session, allowing the experimental toolkit available for use in trial-by-trial tasks for rodents to be applied to the study of planning. We take advantage of one part of this toolkit to address a perennially controversial issue in planning research: the role of the dorsal hippocampus. Although prospective representations in the hippocampus have been proposed to support model-based planning, intact planning in hippocampally damaged animals has been observed in a number of assays. Combining formal algorithmic behavioral analysis with muscimol inactivation, we provide the first causal evidence directly linking dorsal hippocampus with planning behavior. The results reported, and the methods introduced, open the door to new and more detailed investigations of the neural mechanisms of planning, in the hippocampus and throughout the brain.

neuroscience

The successor representation in human reinforcement learning

Theories of reward learning in neuroscience have focused on two families of algorithms, thought to capture deliberative vs. habitual choice. \"Model-based\" algorithms compute the value of candidate actions from scratch, whereas \"model-free\" algorithms make choice more efficient but less flexible by storing pre-computed action values. We examine an intermediate algorithmic family, the successor representation (SR), which balances flexibility and efficiency by storing partially computed action values: predictions about future events. These pre-computation strategies differ in how they update their choices following changes in a task. SRs reliance on stored predictions about future states predicts a unique signature of insensitivity to changes in the tasks sequence of events, but flexible adjustment following changes to rewards. We provide evidence for such differential sensitivity in two behavioral studies with humans. These results suggest that the SR is a computational substrate for semi-flexible choice in humans, introducing a subtler, more cognitive notion of habit.

neuroscience

Predictive representations can link model-based reinforcement learning to model-free mechanisms

Humans and animals are capable of evaluating actions by considering their long-run future rewards through a process described using model-based reinforcement learning (RL) algorithms. The mechanisms by which neural circuits perform the computations prescribed by model-based RL remain largely unknown; however, multiple lines of evidence suggest that neural circuits supporting model-based behavior are structurally homologous to and overlapping with those thought to carry out model-free temporal difference (TD) learning. Here, we lay out a family of approaches by which model-based computation may be built upon a core of TD learning. The foundation of this framework is the successor representation, a predictive state representation that, when combined with TD learning of value predictions, can produce a subset of the behaviors associated with model-based learning, while requiring less decision-time computation than dynamic programming. Using simulations, we delineate the precise behavioral capabilities enabled by evaluating actions using this approach, and compare them to those demonstrated by biological organisms. We then introduce two new algorithms that build upon the successor representation while progressively mitigating its limitations. Because this framework can account for the full range of observed putatively model-based behaviors while still utilizing a core TD framework, we suggest that it represents a neurally plausible family of mechanisms for model-based evaluation.\n\nAuthor SummaryAccording to standard models, when confronted with a choice, animals and humans rely on two separate, distinct processes to come to a decision. One process deliberatively evaluates the consequences of each candidate action and is thought to underlie the ability to flexibly come up with novel plans. The other process gradually increases the propensity to perform behaviors that were previously successful and is thought to underlie automatically executed, habitual reflexes. Although computational principles and animal behavior support this dichotomy, at the neural level, there is little evidence supporting a clean segregation. For instance, although dopamine -- famously implicated in drug addiction and Parkinsons disease -- currently only has a well-defined role in the automatic process, evidence suggests that it also plays a role in the deliberative process. In this work, we present a computational framework for resolving this mismatch. We show that the types of behaviors associated with either process could result from a common learning mechanism applied to different strategies for how populations of neurons could represent candidate actions. In addition to demonstrating that this account can produce the full range of flexible behavior observed in the empirical literature, we suggest experiments that could detect the various approaches within this framework.

neuroscience