bioRxiv · 10.1101/2024.11.01.621621
Policy optimization emerges from noisy representation learning
Abstract
Biological nervous systems learn both internal representations of the world and behavioral policies for acting within it. Motivated by growing evidence that representation learning is a fundamental principle underlying synaptic plasticity, we introduce Neural Stochastic Modulation (NSM): a theory of learning in which policy optimization emerges from reward-modulated noise layered on top of plasticity rules designed for representation learning. In NSM, reward-modulated noise shapes the steady-state weight distribution, guiding the network toward solutions that capture meaningful features while also maximizing reward. Interestingly, the evolving internal representations produced by our model mirror neural coding changes observed experimentally during task learning. Our results suggest that reward-modulated noise can serve as a minimal and biologically plausible mechanism for integrating representation and policy learning in the brain.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Brenner, J. W., Li, C., Kreiman, G.. 2024-11-03. Policy optimization emerges from noisy representation learning. https://doi.org/10.1101/2024.11.01.621621
Cite the original work for its findings. Save a collection to share your selection of sources.