Search bioRxivSearch

Biology subjects

Frank, M. J.

Publications and source records attributed to Frank, M. J..

8 recordsLinked to original sources

Hierarchical Bayesian inference for concurrent model fitting and comparison for group studies

Computational modeling plays an important role in modern neuroscience research. Much previous research has relied on statistical methods, separately, to address two problems that are actually interdependent. First, given a particular computational model, Bayesian hierarchical techniques have been used to estimate individual variation in parameters over a population of subjects, leveraging their population-level distributions. Second, candidate models are themselves compared, and individual variation in the expressed model estimated, according to the fits of the models to each subject. The interdependence between these two problems arises because the relevant population for estimating parameters of a model depends on which other subjects express the model. Here, we propose a hierarchical Bayesian inference (HBI) framework for concurrent model comparison, parameter estimation and inference at the population level, combining previous approaches. We show that this framework has important advantages for both parameter estimation and model comparison theoretically and experimentally. The parameters estimated by the HBI show smaller errors compared to other methods. Model comparison by HBI is robust against outliers and is not biased towards overly simplistic models. Furthermore, the fully Bayesian approach of HBI enables researchers to quantify uncertainty in group parameter estimates, for each candidate model separately, and to perform statistical tests on parameters of a population.

neuroscience

Impaired expected value computations in schizophrenia are associated with a reduced ability to integrate reward probability and magnitude of recent outcomes.

ABSTRACTO_ST_ABSBackgroundC_ST_ABSMotivational deficits in people with schizophrenia (PSZ) are associated with an inability to integrate the magnitude and probability of previous outcomes. The mechanisms that underlie probability-magnitude integration deficits, however, are poorly understood. We hypothesized that increased reliance on \"value-less\" stimulus-response associations, in lieu of expected value (EV)-based learning, could drive probability-magnitude integration deficits in PSZ with motivational deficits.\n\nMethodsHealthy volunteers (n= 38) and PSZ (n=49) completed a reinforcement learning paradigm consisting of four stimulus pairs. Reward magnitude (3/2/1/0 points) and probability (90%/80%/20%/10%) together determined each stimulus EV. Following a learning phase, new and familiar stimulus pairings were presented. Participants were asked to select stimuli with the highest reward value.\n\nResultsPSZ with high motivational deficits made increasingly less optimal choices as the difference in reward value (probability*magnitude) between two competing stimuli increased. Using a previously-validated computational hybrid model, PSZ relied less on EV (\"Q-learning\") and more on stimulus-response learning (\"actor-critic\"), which correlated with SANS motivational deficit severity. PSZ specifically failed to represent reward magnitude, consistent with model demonstrations showing that response tendencies in the actor-critic were preferentially driven by reward probability.\n\nConclusionsProbability-magnitude deficits in PSZ with motivational deficits arise from underutilization of EV in favor of reliance on value-less stimulus-response associations. Consistent with previous work and confirmed by our computational hybrid framework, probability-magnitude integration deficits were driven specifically by a failure to represent reward magnitude. This work reconfirms the importance of decreased Q-learning/increased actor-critic-type learning as an explanatory framework for a range of EV deficits in PSZ.

neuroscience

Positive reward prediction errors strengthen incidental memory encoding

The dopamine system is thought to provide a reward prediction error signal that facilitates reinforcement learning and reward-based choice in corticostriatal circuits. While it is believed that similar prediction error signals are also provided to temporal lobe memory systems, the impact of such signals on episodic memory encoding has not been fully characterized. Here we develop an incidental memory paradigm that allows us to 1) estimate the influence of reward prediction errors on the formation of episodic memories, 2) dissociate this influence from other factors such as surprise and uncertainty, 3) test the degree to which this influence depends on temporal correspondence between prediction error and memoranda presentation, and 4) determine the extent to which this influence is consolidation-dependent. We find that when choosing to gamble for potential rewards during a primary decision making task, people encode incidental memoranda more strongly even though they are not aware that their memory will be subsequently probed. Moreover, this strengthened encoding scales with the reward prediction error, and not overall reward, experienced selectively at the time of memoranda presentation (and not before or after). Finally, this strengthened encoding is identifiable within a few minutes and is not substantially enhanced after twenty-four hours, indicating that it is not consolidation-dependent. These results suggest a computationally and temporally specific role for putative dopaminergic reward prediction error signaling in memory formation.

neuroscience

Impaired expected value computations coupled with overreliance on prediction error learning in schizophrenia

BackgroundWhile many have emphasized impaired reward prediction error (RPE) signaling in schizophrenia, multiple studies suggest that some decision-making deficits may arise from overreliance on RPE systems together with a compromised ability to represent expected value. Guided by computational frameworks, we formulated and tested two scenarios in which maladaptive representation of expected value should be most evident, thereby delineating conditions that may evoke decision-making impairments in schizophrenia.\n\nMethodsIn a modified reinforcement learning paradigm, 42 medicated people with schizophrenia (PSZ) and 36 healthy volunteers learned to select the most frequently rewarded option in a 75-25 pair: once when presented with more deterministic (90-10) and once when presented with more probabilistic (60-40) pairs. Novel and old combinations of choice options were presented in a subsequent transfer phase. Computational modeling was employed to elucidate contributions from RPE systems (\"actor-critic\") and expected value (\"Q-leaming\").\n\nResultsPSZ showed robust performance impairments with increasing value difference between two competing options, which strongly correlated with decreased contributions from expected value-based (\"Q-leaming\") learning. Moreover, a subtle yet consistent contextual choice bias for the \"probabilistic\" 75 option was present in PSZ, which could be accounted for by a context-dependent RPE in the \"actor-critic\".\n\nConclusionsWe provide evidence that decision-making impairments in schizophrenia increase monotonically with demands placed on expected value computations. A contextual choice bias is consistent with overreliance on RPE-based learning, which may signify a deficit secondary to the maladaptive representation of expected value. These results shed new light on conditions under which decisionmaking impairments may arise.

neuroscience

A control theoretic model of adaptive behavior in dynamic environments

To behave adaptively in environments that are noisy and non-stationary, humans and other animals must monitor feedback from their environment and adjust their predictions and actions accordingly. An under-studied approach for modeling these adaptive processes comes from the engineering field of control theory, which provides general principles for regulating dynamical systems, often without requiring a generative model. The proportional-integral-derivative (PID) controller is one of the most popular models of industrial process control. The proportional term is analogous to the \"delta rule\" in psychology, adjusting estimates in proportion to each successive error in prediction. The integral and derivative terms augment this update to simultaneously improve accuracy and stability. Here, we tested whether the PID algorithm can describe how people sequentially adjust their predictions in response to new information. Across three experiments, we found that the PID controller was an effective model of participants decisions in noisy, changing environments. In Experiment 1, we re-analyzed a change-point detection experiment, and showed that participants behavior incorporated elements of PID updating. In Experiments 2-3 we developed a task with gradual transitions that we optimized to detect PID-like adjustments. In both experiments, the PID model offered better descriptions of behavioral adjustments than both the classical delta-rule model and its more sophisticated variant, the Kalman filter. We further examined how participants weighted different PID terms in response to salient environmental events, finding that these control terms were modulated by reward, surprise, and outcome entropy. These experiments provide preliminary evidence that adaptive behavior in dynamic environments resembles PID control.

neuroscience

Cross-task contributions of fronto-basal ganglia circuitry in response inhibition and conflict-induced slowing

Why are we so slow in choosing the lesser of two evils? We considered whether such slowing relates to uncertainty about the value of these options, which arises from the tendency to avoid them during learning, and whether such slowing relates to fronto-subthalamic inhibitory control mechanisms. 49 participants performed a reinforcement-learning task and a stop-signal task while fMRI was recorded. A reinforcement-learning model was used to quantify learning strategies. Individual differences in lose-lose slowing related to information uncertainty due to sampling, and independently, to less efficient response inhibition in the stop-signal task. Neuroimaging analysis revealed an analogous dissociation: subthalamic nucleus (STN) BOLD activity related to variability in stopping latencies, whereas weaker fronto-subthalamic connectivity related to slowing and information sampling. Across tasks, fast inhibitors increased STN activity for successfully cancelled responses in the stop task, but decreased activity for lose-lose choices. These data support the notion that fronto-STN communication implements a rapid but transient brake on response execution, and that slowing due to decision uncertainty could result from an inefficient release of this \"hold your horses\" mechanism.

neuroscience

Compositional clustering in task structure learning

Humans are remarkably adept at generalizing knowledge between experiences in a way that can be difficult for computers. Often, this entails generalizing constituent pieces of experiences that do not fully overlap, but nonetheless share useful similarities with, previously acquired knowledge. However, it is often unclear how knowledge gained in one context should generalize to another. Previous computational models and data suggest that rather than learning about each individual context, humans build latent abstract structures and learn to link these structures to arbitrary contexts, facilitating generalization. In these models, task structures that are more popular across contexts are more likely to be revisited in new contexts. However, these models predict that structures are either re - used as a whole or created from scratch, prohibiting the ability to generalize constituent parts of learned structures. This contrasts with ecological settings, where some aspects of task structure, such as the transition function, will be shared between context separately from other aspects, such as the reward function. Here, we develop a novel non - parametric Bayesian agent that forms independent latent clusters for transition and reward functions that may have different popularity across contexts. We compare this agent to an agent that jointly clusters both across a range of task domains. We show that relative performance of the two agents depends on the statistics of the task domain, including the mutual information between transition and reward functions in the environment, and the stochasticity of the observations. We formalize our analysis through an information theoretic account of the priors, and develop a meta learning agent that can dynamically arbitrate between strategies across task domains. We argue that this provides a first step in allowing for compositional structures in reinforcement learners, which should be provide a better model of human learning and additional flexibility for artificial agents.\n\nAuthor summaryA musician may learn to generalize behaviors across instruments for different purposes, for example, reusing hand motions used when playing classical on the flute to play jazz on the saxophone. Conversely, she may learn to play a single song across many instruments that require completely distinct physical motions, but nonetheless transfer knowledge between them. This degree of compositionality is often absent from computational frameworks of learning, forcing agents either to generalize entire learned policies or to learn new policies from scratch. Here, we propose a solution to this problem that allows an agent to generalize components of a policy independently and compare it to an agent that generalizes components as a whole. We show that the degree to which one form of generalization is favored over the other is dependent on the features of task domain, with independent generalization of task components favored in environments with weak relationships between components or high degrees of noise and joint generalization of task components favored when there is a clear, discoverable relationship between task components. Furthermore, we show that the overall meta structure of the environment can be learned and leveraged by an agent that dynamically arbitrates between these forms of structure learning.

neuroscience

Chunking as a rational strategy for lossy data compression in visual working memory tasks.

The nature of capacity limits for visual working memory has been the subject of an intense debate that has relied on models that assume items are encoded independently. Here we propose that instead, similar features are jointly encoded through a \"chunking\" process to optimize performance on visual working memory tasks. We show that such chunking can: 1) facilitate performance improvements for abstract capacity-limited systems, 2) be optimized through reinforcement, 3) be implemented by center-surround dynamics, and 4) increase effective storage capacity at the expense of recall precision. Human performance on a variant of a canonical working memory task demonstrated performance advantages, precision detriments, inter-item dependencies, and trial-to-trial behavioral adjustments diagnostic of performance optimization through center-surround chunking. Models incorporating center-surround chunking provided a better quantitative description of human performance in our study as well as in a meta-analytic dataset, and apparent differences in working memory capacity across individuals were attributable to individual differences in the implementation of chunking. Our results reveal a normative rationale for center-surround connectivity in working memory circuitry, call for re-evaluation of memory performance differences that have previously been attributed to differences in capacity, and support a more nuanced view of visual working memory capacity limitations: strategic tradeoff between storage capacity and memory precision through chunking contribute to flexible capacity limitations that include both discrete and continuous aspects.

neuroscience