Search bioRxivSearch

Biology subjects

Farashahi, S.

Publications and source records attributed to Farashahi, S..

4 recordsLinked to original sources

On the flexibility of basic risk attitudes in monkeys

Monkeys and other animals appear to share with humans two risk attitudes predicted by prospect theory: an inverse-S-shaped probability weighting function and a steeper utility curve for losses than for gains. These findings suggest that such preferences are stable traits with common neural substrates. We hypothesized instead that animals tailor their preferences to subtle changes in task contexts, making risk attitudes flexible. Previous studies used a limited number of outcomes, trial types, and contexts. To gain a broader perspective, we examined two large datasets of male macaques risky choices: one from a task with real (juice) gains and another from a token task with gains and losses. In contrast to previous findings, monkeys were risk-seeking for both gains and losses (i.e. lacked a reflection effect) and showed steeper gain than loss curves (loss-seeking). Utility curves for gains were substantially different in the two tasks. Monkeys showed nearly linear probability weightings in one task and S-shaped ones in the other; neither task produced a consistent inverse-S-shaped curve. To account for these observations, we developed and tested various computational models of the processes involved in the construction of reward value. We found that adaptive differential weighting of prospective gamble outcomes could partially account for the observed differences in the utility functions across the two experiments and thus, provide a plausible mechanism underlying flexible risk attitudes. Together, our results support the idea that risky choices are flexibly constructed at the time of elicitation and place important constraints on neural models of economic choice.

neuroscience

Dynamic combination of sensory and reward information under time pressure

When making choices, collecting more information is beneficial but comes at the cost of sacrificing time that could be allocated to making other potentially rewarding decisions. To investigate how the brain balances these costs and benefits, we conducted a series of novel experiments in humans and simulated various computational models. Under six levels of time pressure, subjects made decisions either by integrating sensory information over time or by dynamically combining sensory and reward information over time. We found that during sensory integration, time pressure reduced performance as the deadline approached, and choice was more strongly influenced by the most recent sensory evidence. By fitting performance and reaction time with various models we found that our experimental results are more compatible with leaky integration of sensory information with an urgency signal or a decision process based on stochastic transitions between discrete states modulated by an urgency signal. When combining sensory and reward information, subjects spent less time on integration than optimally prescribed when reward decreased slowly over time, and the most recent evidence did not have the maximal influence on choice. The suboptimal pattern of reaction time was partially mitigated in an equivalent control experiment in which sensory integration over time was not required, indicating that the suboptimal response time was influenced by the perception of imperfect sensory integration. Meanwhile, during combination of sensory and reward information, performance did not drop as the deadline approached, and response time was not different between correct and incorrect trials. These results indicate a decision process different from what is involved in the integration of sensory information over time. Together, our results not only reveal limitations in sensory integration over time but also illustrate how these limitations influence dynamic combination of sensory and reward information.

neuroscience

Influence of learning strategy on response time during complex value-based learning and choice

Measurements of response time (RT) have long been used to infer neural processes underlying various cognitive functions such as working memory, attention, and decision making. However, it is currently unknown if RT is also informative about various stages of value-based choice, particularly how reward values are constructed. To investigate these questions, we analyzed the pattern of RT during a set of multi-dimensional learning and decision-making tasks that can prompt subjects to adopt different learning strategies. In our experiments, subjects could use reward feedback to directly learn reward values associated with possible choice options (object-based learning). Alternatively, they could learn reward values of options features (e.g. color, shape) and combine these values to estimate reward values for individual options (feature-based learning). We found that RT was slower when the difference between subjects estimates of reward probabilities for the two alternative objects on a given trial was smaller. Moreover, RT was overall faster when the preceding trial was rewarded or when the previously selected object was present. These effects, however, were mediated by an interaction between these factors such that subjects were faster when the previously selected object was present rather than absent but only after unrewarded trials. Finally, RT reflected the learning strategy (i.e. object-based or feature-based approach) adopted by the subject on a trial-by-trial basis, indicating an overall faster construction of reward value and/or value comparison during object-based learning. Altogether, these results demonstrate that the pattern of RT can be informative about how reward values are learned and constructed during complex value-based learning and decision making.

neuroscience

Your favorite color makes learning more adaptable and precise

Learning from reward feedback is essential for survival but can become extremely challenging with myriad choice options. Here, we propose that learning reward values of individual features can provide a heuristic for estimating reward values of choice options in dynamic, multidimensional environments. We hypothesized that this feature-based learning occurs not just because it can reduce dimensionality, but more importantly because it can increase adaptability without compromising precision of learning. We experimentally tested this hypothesis and found that in dynamic environments, human subjects adopted feature-based learning even when this approach does not reduce dimensionality. Even in static, low-dimensional environments, subjects initially adopted feature-based learning and gradually switched to learning reward values of individual options, depending on how accurately objects values can be predicted by combining feature values. Our computational models reproduced these results and highlight the importance of neurons coding feature values for parallel learning of values for features and objects.

neuroscience