Search bioRxivSearch

Biology subjects

Balleine, B. W.

Publications and source records attributed to Balleine, B. W..

5 recordsLinked to original sources

Integrated accounts of behavioral and neuroimaging data using flexible recurrent neural network models

Neuroscience studies of human decision-making abilities commonly involve sub-jects completing a decision-making task while BOLD signals are recorded using fMRI. Hypotheses are tested about which brain regions mediate the effect of past experience, such as rewards, on future actions. One standard approach to this is model-based fMRI data analysis, in which a model is fitted to the behavioral data, i.e., a subjects choices, and then the neural data are parsed to find brain regions whose BOLD signals are related to the models internal signals. However, the internal mechanics of such purely behavioral models are not constrained by the neural data, and therefore might miss or mischaracterize aspects of the brain. To address this limitation, we introduce a new method using recurrent neural network models that are flexible enough to be jointly fitted to the behavioral and neural data. We trained a model so that its internal states were suitably related to neural activity during the task, while at the same time its output predicted the next action a subject would execute. We then used the fitted model to create a novel visualization of the relationship between the activity in brain regions at different times following a reward and the choices the subject subsequently made. Finally, we validated our method using a previously published dataset. We found that the model was able to recover the underlying neural substrates that were discovered by explicit model engineering in the previous work, and also derived new results regarding the temporal pattern of brain activity.

neuroscience

Models that learn how humans learn: the case of depression and bipolar disorders

Computational models of learning and decision-making processes in the brain play an important role in many domains. Such models typically have a constrained structure and make specific assumptions about the underlying human learning processes; these may make them underfit observed behaviours. Here we suggest an alternative method based on learning-to-learn approaches, using recurrent neural networks (RNNs) as a flexible family of models that have sufficient capacity to represent the complex learning and decision-making strategies used by humans. In this approach, an RNN is trained to predict the next action that a subject will take in a decision-making task, and in this way, learns to imitate the processes underlying subjects choices and their learning abilities. We demonstrate the benefits of this approach with a new dataset containing behaviour of uni-polar depression (n=34), bipolar (n=33) and control (n=34) participants in a two-armed bandit task. The results indicate that the new approach is better than baseline reinforcement-learning methods in terms of overall performance and its capacity to predict subjects choices. We show that the model can be interpreted using off-policy simulations, and thereby provide a novel clustering of subjects learning processes - something that often eludes traditional approaches to modelling and behavioural analysis.

bioinformatics

Learning the structure of the world: The adaptive nature of state-space and action representations in multi-stage decision-making

State-space and action representations form the building blocks of decision-making processes in the brain; states map external cues to the current situation of the agent whereas actions provide the set of motor commands from which the agent can choose to achieve specific goals. Although these factors differ across environments, it is not currently known whether or how accurately state and action representations are acquired by the agent because previous experiments have typically provided this information a priori through instruction or pre-training. Here we show that, in the absence of such a priori knowledge, state and action representations adapt to reflect the structure of the world. We used a sequential decision-making task in rats in which they were required to pass through multiple states before reaching the goal, and for which the number of states and how they map onto external cues were not known a priori. We found that, early in training, animals selected actions as if the task was not sequential and outcomes were the immediate consequence of the most proximal action. During the course of training, however, rats recovered the true structure of the environment and made decisions based on the expanded state-space, reflecting the multiple stages of the task. We found a similar pattern with actions; early in training animals only considered the execution of single actions whereas, after training, they created useful action sequences that expanded the set of available actions. We conclude that the profile of choices shows a gradual shift from simple representations of actions and states to more complex structures compatible with the structure of the world.

animal behavior and cognition

The Algorithmic Neuroanatomy Of Action-Outcome Learning

Although it is well known that animals can encode the consequences of their actions and can use this information to control action selection and evaluation, it is not known what learning rules control action-outcome (AO) learning. Here we trained participants to encode specific AO associations whilst undergoing functional imaging (fMRI) and used computational modelling to evaluate competing models. This analysis revealed that a Kalman filter, which learned the unique causal effect of each action, best characterized AO learning and found the medial prefrontal cortex differentiated the unique effect of actions from background effects. We subsequently extended these findings to show that mPFC participated in a circuit with parietal cortex and caudate nucleus to segregate distinct contributions to AO learning. The results extend our understanding of goal-directed learning and demonstrate that sensitivity to the causal relationship between actions and outcomes guides goal-directed learning rather than contiguous state-action relations.

neuroscience

Optimal Response Vigor and Choice Under Non-stationary Outcome Values

Within a rational framework a decision-maker selects actions based on the reward-maximisation principle which stipulates they acquire outcomes with the highest values at the lowest cost. Action selection can be divided into two dimensions: selecting an action from several alternatives, and choosing its vigor, i.e., how fast the selected action should be executed. Both of these dimensions depend on the values of the outcomes, and these values are often affected as more outcomes are consumed, and so are the actions. Despite this, previous works have addressed the computational substrates of optimal actions only in the specific condition that the values of outcomes are constant, and it is still unknown what the optimal actions are when the values of outcomes are non-stationary. Here, based on an optimal control framework, we derive a computational model for optimal actions under non-stationary outcome values. The results imply that even when the values of outcomes are changing, the optimal response rate is constant rather than decreasing. This finding shows that, in contrast to previous theories, the commonly observed changes in the actions cannot be purely attributed to the changes in the outcome values. We then prove that this observation can be explained based on the uncertainty about temporal horizons; e.g., in the case of experimental protocols, the session duration. We further show that when multiple outcomes are available, the model explains probability matching as well as maximisation choice strategies. The model provides, therefore, a quantitative analysis of optimal actions and explicit predictions for future testing.

animal behavior and cognition