Search bioRxiv⌕ Search

Biology subjects

Greenstreet, F.

Publications and source records attributed to Greenstreet, F..

2 recordsLinked to original sources

Why motor learning involves multiple systems: an algorithmic perspective

The initial stage of learning motor skills involves exploring vast action spaces, making it impractical to learn the value of every possible action independently. This poses a challenge for standard reinforcement learning approaches, which excel in constrained domains but struggle when the space of possible actions is high-dimensional. Recent work in machine learning has sought to mitigate this problem by combining deep re-inforcement learning with a supervised learning system that reduces the complexity of the control policy space by learning low-dimensional embeddings of an action space. Here, we propose that in the mammalian brain, the cortico-cerebellar network learns these low-dimensional action embeddings in a supervised way, while the basal ganglia learn value and policies in this action embedding space using reinforcement learning. We trained this model on reaching tasks and show that, contrary to traditional models of the basal ganglia, it recapitulates features of neural activity whereby similar reaching movements are associated with similar neural activity patterns in the basal ganglia. We also demonstrate a link between learning these low-dimensional action embeddings and both generalisation and the limits of multi-task adaptation in human behavioural studies. Through this framework, we propose a novel computational view of how key motor regions of the brain interact to efficiently learn a new skill.

neuroscience↗

Action prediction error: a value-free dopaminergic teaching signal that drives stable learning

Animals choice behavior is characterized by two main tendencies: taking actions that led to rewards and repeating past actions. Theory suggests these strategies may be reinforced by different types of dopaminergic teaching signals: reward prediction error (RPE) to reinforce value-based associations and movement-based action prediction errors to reinforce value-free repetitive associations. Here we use an auditory-discrimination task in mice to show that movement-related dopamine activity in the tail of the striatum encodes the hypothesized action prediction error signal. Causal manipulations reveal that this prediction error serves as a value-free teaching signal that supports learning by reinforcing repeated associations. Computational modeling and experiments demonstrate that action prediction errors alone cannot support reward-guided learning but when paired with the RPE circuitry they serve to consolidate stable sound-action associations in a value-free manner. Together we show that there are two types of dopaminergic prediction errors that work in tandem to support learning.

neuroscience↗