Search bioRxivSearch

Biology subjects

Verstynen, T. V.

Publications and source records attributed to Verstynen, T. V..

2 recordsLinked to original sources

A way around the exploration-exploitation dilemma

Balancing exploration with exploitation is seen as a mathematically intractable dilemma that all animals face. In this paper, we provide an alternative view of this classic problem that does not depend on exploring to optimize for reward. We argue that the goal of exploration should be pure curiosity, or learning for learnings sake. Through theory and simulations we prove that explore-exploit problems based on this can be solved by a simple rule that yields optimal solutions: when information is more valuable than rewards, be curious, otherwise seek rewards. We show that this rule performs well and robustly under naturalistic constraints. We suggest three criteria can be used to distinguish our approach from other theories.

animal behavior and cognition

Corticostriatal synaptic weight evolution in a two-alternative forced choice task

In natural environments, mammals can efficiently select actions based on noisy sensory signals and quickly adapt to unexpected outcomes to better exploit opportunities that arise in the future. Such feedback-based changes in behavior rely on long term plasticity within cortico-basal-ganglia-thalamic networks, driven by dopaminergic modulation of cortical inputs to the direct and indirect pathway neurons of the striatum. While the firing rates of corticostriatal neurons have been shown to adapt across a range of feedback conditions, it remains difficult to directly assess the corticostriatal synaptic weight changes that contribute to these adaptive firing rates. In this work, we simulate a computational model for the evolution of corticostriatal synaptic weights based on a spike timing-dependent plasticity rule driven by dopamine signaling that is induced by outcomes of actions in the context of a two-alternative forced choice task. Results show that plasticity predominantly impacts direct pathway weights, which evolve to drive action selection toward a more-rewarded action in settings with deterministic reward outcomes. After the model is tuned based on such fixed reward scenarios, its performance agrees with the results of behavioral experiments carried out with probabilistic reward paradigms.

animal behavior and cognition