bioRxiv · 10.1101/2024.10.12.618033
A computational principle of habit formation
Abstract
1Actions are influenced by multiple decision-making systems - including a goal-directed system that favors rewarded actions and a habit system that repeats past actions - but precisely when one system prevails is not known. We show that when competition between these systems is resolved by a winner-take-all mechanism, the precise condition for the emergence of habits can be cast in terms of the well-known probability matching principle. The theory embodies a trade-off in which exploitation, or overmatching, maximizes reward but strengthens habits, while paradoxically, exploration preserves goal-directed behavior by sacrificing rewards. This tradeoff can be averted if learning operates on abstract latent state representations whereby knowing the broader context allows for switching between two habits instead of avoiding one, thus maximizing rewards without forfeiting flexibility. The theory explains a range of animal behaviors as well as task-dependent effects of striatal manipulation, and suggests that neural mechanisms governing exploration implicitly control arbitration between decision-making systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lakshminarasimhan, K. J.. 2024-10-13. A computational principle of habit formation. https://doi.org/10.1101/2024.10.12.618033
Cite the original work for its findings. Save a collection to share your selection of sources.