Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.04.10.717847

Modulation of feature attention by reward prediction error explains value learning behavior

Abstract

Adaptive behavior requires learning the value of environmental features while selectively attending to those most likely to yield reward. Reward prediction errors (RPEs) drive value learning and learned values guide attention, yet the computational function linking RPEs to attentional modulation remains unspecified. Here, we developed a reinforcement learning model with a perceptual front-end to investigate how value and RPE signals modulate attentional gain during learning. We compared five candidate RPE-attention transfer functions, each combined with either single- or multi-focus attention, against behavioral data from two adult male rhesus macaques performing a color-value learning task with shifting reward contingencies. Monkeys exhibited rapid initial learning followed by sub-optimal asymptotic accuracy. Overall, single-focus architectures consistently outperformed multi-focus counterparts on matching monkey errors, indicating that macaques collapse the value distribution into a winner-take-all attentional focus. Furthermore, the "Switch" model, in which attention targets the highest-valued feature but transiently inverts following negative RPEs, produced the fastest exploration dynamics following target switches and, together with the Absolute Value model, yielded decision confidence trajectories that positively correlated with empirical reaction times. In support of this, single-neuron correlation analyses revealed that 27 - 42% of neurons in prefrontal cortex, frontal eye fields, and lateral intraparietal area encoded previous-trial RPE at the time of next trial onset. In total, we conclude that capacity constrained attention that inverts its focus after negative RPE best explains value learning dynamics. These results provide a normative account for why biological learners sacrifice asymptotic precision for rapid adaptation in volatile environments. Significance StatementLearning which features of the environment predict reward requires both reinforcement learning and selective attention, yet how these processes interact algorithmically remains unknown. We developed a computational model that specifies how reward prediction errors dynamically adjust the gain of feature-based attention during value learning. Testing competing hypotheses against macaque behavioral data, we show that a "Switch" mechanism, in which negative prediction errors transiently invert attentional focus away from the highest-valued feature, best captures primate learning dynamics. This architecture reveals a principled trade-off: the brain sacrifices asymptotic accuracy for rapid detection of environmental change, using error-triggered attentional inversion as a directed exploration strategy. These findings bridge reinforcement learning theory and attention research by identifying the transfer function linking prediction errors to sensory gain modulation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Leukos, M. L., Liang, A., Lindsay, G. W.. 2026-04-11. Modulation of feature attention by reward prediction error explains value learning behavior. https://doi.org/10.64898/2026.04.10.717847

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Functional validation of allele-specific LMNB1 silencing in patient-derived astrocytes as a therapeutic option for Autosomal Dominant Leukodystrophy

Adult-onset Autosomal Dominant Leukodystrophy (ADLD) is a rare fatal leukodystrophy caused by increased LMNB1 gene dosage, most commonly resulting from duplication of the LMNB1 locus. Because ADLD is a gene dosage disorder, selective reduction of pathological LMNB1 expression represents a rational therapeutic strategy. Although allele-specific RNA interference has previously been shown to lower LMNB1 levels in patient-derived fibroblasts and directly reprogrammed neurons, its therapeutic effects have not been evaluated in disease-relevant human glial cells or using functional efficacy endpoints. Here, we established human induced pluripotent stem cell-derived astrocytes from ADLD patients as a human glial model in which to validate allele-specific LMNB1 silencing across molecular, cellular, and functional readouts. ADLD astrocytes recapitulated increased LMNB1 expression and characteristic nuclear abnormalities and displayed transcriptional alterations affecting extracellular matrix organization, calcium homeostasis, metabolism and RNA processing. Functionally, these cells also exhibited functional phenotypes suitable for therapeutic evaluation: astrocyte-conditioned medium impaired the viability of both murine and human oligodendroglial cultures, while conditioned-medium and direct astrocyte-seeding paradigms revealed impaired post-lesion myelin recovery in lysolecithin-treated cerebellar organotypic slices. Allele-specific LMNB1 silencing restored physiological LMNB1 levels, corrected nuclear abnormalities, attenuated astrocyte-mediated oligodendroglial toxicity, improved post-lesion myelin recovery, and was associated with selective transcriptional programs associated with extracellular support and cholesterol metabolism. Together, these findings provide molecular, cellular, and functional validation of allele-specific LMNB1 dosage correction in patient-derived human astrocytes and offer key support for LMNB1-lowering strategies in disease-relevant human glial cells.

neuroscience↗

Perceptual integration of multisensory haptic, visual, and auditory feedback for roughness discrimination in augmented reality

Understanding how our different senses interact to shape our perception is essential to design realistic and immersive virtual and augmented reality (VR/AR) experiences. The present study investigated how roughness perception can be modulated through haptic, visual, and auditory cues in AR using a vibrotactile wristband. Participants compared virtual textures varying in vibration frequency/amplitude, visual grain size, and friction sound. Results revealed strong linear relationships between stimulus parameters and perceived roughness, with haptic frequency and visual cues driving the highest discrimination performance. Adding non-informative sensory feedback reduced perceptual sensitivity, acting as noise. Individual differences emerged: participants who rated haptic as the easiest modality showed greater sensitivity to haptic variations, while visual-reliant participants performed better with visual cues. We conclude that roughness in AR can be systematically manipulated, but is vulnerable to perceptual interference from irrelevant inputs, where our work provides actionable insights for implementing optimized and adaptive AR/VR interfaces.

neuroscience↗

Structural and functional MRI signatures of Gambling Disorder: a case-control study

Gambling disorder (GD) is a behavioural addiction that may help identify addiction-related neural features without the direct neurobiological effects of a primary substance of dependence. We examined regional grey matter volume (GMV) and resting-state functional connectivity (rsFC) in the same well-characterised sample. Eighteen men with GD and 21 matched healthy controls underwent high-resolution structural and resting-state functional MRI. GMV was quantified across 214 cortical and subcortical regions, and seed-based rsFC analyses focused on striatal subdivisions and mesocorticolimbic regions. Group differences were evaluated using permutation testing and cluster-corrected mixed-effects modelling. GD was associated with lower GMV in the ventromedial prefrontal cortex, orbitofrontal regions and other cortical and subcortical areas, alongside higher GMV in a subset of limbic and default-mode regions. Participants with GD also showed lower connectivity between the limbic striatum and the hippocampus, thalamus and putamen. In exploratory analyses, somatomotor connectivity was positively associated with gambling severity (Problem Gambling Severity Index: Spearman's rho = 0.71, p = 0.003, false-discovery-rate-adjusted q = 0.016). Structural and functional findings overlapped spatially in regions associated with valuation, memory, reward and habit formation, but regional GMV did not mediate group differences in rsFC. These findings are broadly consistent with corticostriatal models of GD and identify candidate circuit-level differences for independent replication. Larger, more diverse and longitudinal samples are required to establish their reproducibility, temporal direction and clinical relevance.

neuroscience↗