Search bioRxiv⌕ Search

Biology subjects

Brenner, J. W.

Publications and source records attributed to Brenner, J. W..

2 recordsLinked to original sources

Policy optimization emerges from noisy representation learning

Biological nervous systems learn both internal representations of the world and behavioral policies for acting within it. Motivated by growing evidence that representation learning is a fundamental principle underlying synaptic plasticity, we introduce Neural Stochastic Modulation (NSM): a theory of learning in which policy optimization emerges from reward-modulated noise layered on top of plasticity rules designed for representation learning. In NSM, reward-modulated noise shapes the steady-state weight distribution, guiding the network toward solutions that capture meaningful features while also maximizing reward. Interestingly, the evolving internal representations produced by our model mirror neural coding changes observed experimentally during task learning. Our results suggest that reward-modulated noise can serve as a minimal and biologically plausible mechanism for integrating representation and policy learning in the brain.

neuroscience↗

Neuron-level prediction and noise can implement flexible reward-seeking behavior

We show that neural networks can implement reward-seeking behavior using only local predictive updates and internal noise. These networks are capable of autonomous interaction with an environment and can switch between explore and exploit behavior, which we show is governed by attractor dynamics. Networks can adapt to changes in their architectures, environments, or motor interfaces without any external control signals. When networks have a choice between different tasks, they can form preferences that depend on patterns of noise and initialization, and we show that these preferences can be biased by network architectures or by changing learning rates. Our algorithm presents a flexible, biologically plausible way of interacting with environments without requiring an explicit environmental reward function, allowing for behavior that is both highly adaptable and autonomous. Code is available at https://github.com/ccli3896/PaN.

neuroscience↗