Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.06.15.732293

Nonlinear influence of reward volatility on arbitration between multiple learning strategies reflects cost-benefit optimization

Abstract

Action selection involves two systems: a model-free reinforcement learning strategy, which relies on experience with action-outcome pairs, and a model-based reinforcement learning strategy, which enables more flexible behavior via inference using a model of the invariant environmental structure. Although environmental change requires more flexible behavior, the ability of volatility, a higher-order statistic that captures how rapidly or frequently the environment changes, to systematically modulate these strategies remains unclear. We examined the effects of reward volatility on arbitration between model-free and model-based reinforcement learning strategies using two modified two-step decision tasks. In Experiment 1, participants performed tasks with different levels of reward volatility and time pressure. In Experiment 2, we systematically manipulated reward volatility across a broader range to assess the relationship between volatility and learning strategy. Behavioral data were analyzed using model-agnostic one-trial and multitrial back analyses, reinforcement learning simulations, and hierarchical Bayesian model fitting. Across experiments, reward volatility exerted an inverse U-shaped nonlinear effect on the arbitration between model-free and model-based reinforcement learning strategies, as the model-based learning strategy was strongly driven at intermediate levels of reward volatility. These modulation effects were observed only in individuals who had learned the transition structure in the task, whereas those who had not learned the transition structure relied on the model-free learning strategy regardless of reward volatility. Reinforcement learning simulations revealed that the relative advantage of the model-based learning strategy over the model-free learning strategy peaked at intermediate levels of reward volatility. Additionally, increased time pressure shifted behavior toward the model-free learning strategy. These results demonstrated that, humans do not always use the model-based reinforcement learning strategy in uncertain and dynamic environments, even when they are aware of the task structure, supporting cost-benefit optimization. Author SummaryThe ability to flexibly guide behavior by carefully considering future consequences is fundamental to a prominent property of human intelligence and rationality. However, what drives this deliberative system? In this study, we investigated the factors that promote deliberative versus habitual behavior using decision-making tasks with uncertain structures and changing rewards. We found that participants who spontaneously learned the hidden transition structure in the task used this knowledge to guide deliberative behavior. Conversely, participants who did not learn the structure relied primarily on habitual strategies, repeating actions that had previously been rewarded. Among participants who learned the structure, the degree of deliberative behavior changed nonlinearly with reward volatility, in which the speed at which rewards changed over time. We also observed that limiting the decision time reduced deliberative behavior and promoted habitual responding. These findings suggest that under uncertain and dynamic environments, deliberative control is adaptively regulated according to cost-benefit optimization. Our results contribute to understanding how humans flexibly adjust their behavioral control systems in response to environmental conditions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yamada, T., Samejima, K.. 2026-06-19. Nonlinear influence of reward volatility on arbitration between multiple learning strategies reflects cost-benefit optimization. https://doi.org/10.64898/2026.06.15.732293

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

roostR: An R package to examine diel activity patterns from Motus radio telemetry data

1. Monitoring the diel activity patterns of free-living animals is a methodological challenge. Signal strength fluctuations from Very High Frequency (VHF) radio transmitters deployed within the Motus Wildlife Tracking System can be used as a proxy for activity. However, analytical tools to extract behavioral metrics that quantify activity patterns from these data are needed. 2. We developed roostR, an open-source R package that converts Motus detection data into quantitative behavioral metrics, including roost initiation and departure, roost duration, observation time, and restlessness. The package uses a sequential pipeline built around signal volatility and rolling medians to detect transitions between active and inactive states. Default parameter values were tuned using data from 55 dark-eyed juncos (Junco hyemalis) overwintering in southeastern Ohio. 3. We provide an example from a dark-eyed junco over a 58-day period. roostR estimated roost onset on 48 nights and departure on 56 mornings, with higher rolling median signal differences during the day than at night, consistent with a diurnal animal, and variable restlessness periods each night. We also used the package to estimate roost behavior of an American tree sparrow (Spizelloides arborea) over 43-nights. 4. roostR enables researchers to extract individual activity data from Motus datasets. Because all thresholds are user-adjustable, the pipeline is adaptable across species, tag specifications, and ecological contexts, enabling researchers to test hypotheses about how environmental factors influence diel activity patterns.

animal behavior and cognition↗

Characterizing rhythmic wheel-turning behavioral patterns in cockroaches Rhyparobia maderae using machine learning

Organisms must adapt to environmental changes occurring across multiple time scales, with endogenous multiscale clocks coordinating physiology and behavior with recurring environmental rhythms, including the dominant 24-hour cycle and faster ultradian rhythms. The Madeira cockroach (Rhyparobia maderae) provides a suitable model for investigating such multiscale temporal organization. Here, locomotor activity was recorded in running-wheel experiments under constant darkness. While the endogenous circadian clock produces a clearly visible 24-hour rhythm, it remains unknown whether locomotor behavior also exhibits temporal patterns at additional time scales. These temporal patterns cannot be found by classical frequency analysis, as they are veiled by higher harmonics of the circadian rhythm which are in the same frequency range. Unsupervised machine learning methods such as K-Means clustering, self-organizing maps and Gaussian mixture are used in search for fast ultradian rhythms possibly linked to circadian cycles in locomotor activity. Prior to applying these methods, data metrics are defined which characterize bouts of activity (called activity impulses) compared to periods of reduced activity. A stochastic pattern was found in these activity metrics which characterizes the time distance between activity impulses. Across all approaches, a consistent ultradian rhythm of approximately one hour was identified in the timing of the activity maxima. This rhythm was mainly detected during the subjective night, suggesting circadian control, and appears to consist of two components with periods of approximately 40 minutes and 1.5 hours. The method proposed in this paper is applied to two cockroach groups with different levels of activity, and is generalizable to diverse datasets occurring in the form of a time series with a dominant rhythm.

animal behavior and cognition↗

Social inequity aversion and fairness preference in polar bears

Decision-making is key to survival, with choices and behaviour typically shaped by evolutionary pressures to enhance fitness and minimize loss. Inequity aversion, the tendency to respond negatively to unequal reward distributions, is therefore unexpected, as acts of fairness may be costly to the individual in the short term. In social, cooperative species, however, fairness may enhance long-term benefits through stable social interactions. Inequity aversion is believed to promote cooperative social structures and has only been demonstrated in certain social species. Yet, solitary species have not been tested, and whether this behaviour depends on social cooperative structures is unknown. Examining solitary species could clarify whether inequity aversion reflects an adaptation to cooperation or a general cognitive capacity linked to social comparison and reward evaluation. We studied inequity aversion in polar bears (Ursus maritimus), a largely solitary species, using three complementary, non-invasive cognition tests on seven bears from two zoos in the Netherlands. (1) An impunity experiment assessing effort-based inequity, in which bears performed a standardized action to obtain food rewards. (2) A choice-based experiment testing active selection between equal and unequal reward distributions. (3) A no-task control experiment assessing responses to unequal reward distribution without effort. Under effort-based inequity, bears reduced task performance and increased task latency when a partner received a superior reward. Unequal reward distributions did not affect reward acceptance independent of effort. In the choice-based paradigm, bears preferentially selected equal reward distributions with a partner present. Although generalizability to wild populations is limited, these findings demonstrate behavioural aversion of social inequity in a non-cooperative species. Inequality aversion may therefore not be restricted to social, cooperative animals but instead reflect a broadly distributed cognitive mechanism that emerges under relevant social conditions. These findings challenge current theories and advance our understanding of the cognitive foundations of fairness-related behaviour in animals.

animal behavior and cognition↗