Search bioRxiv⌕ Search

Biology subjects

Barfuss, W.

Publications and source records attributed to Barfuss, W..

2 recordsLinked to original sources

Deterministic dynamics of distributional multi-agent reinforcement learning

Understanding how cognition shapes behavior across contexts remains a fundamental challenge for many disciplines. In particular, for the optimism heuristic-i.e., the tendency to overweight positive (relative to negative) information-knowledge remains fragmented, with models developed in specific domains in isolation. Here, we present a unifying computational framework by deriving the deterministic dynamics of distributional multi-agent reinforcement learning. Our approach discretizes reward distributions through a finite set of neurons, consistent with recent empirical findings on distributional coding in the brain. We validate our framework by reproducing established results across three iconic domains spanning individual exploration versus exploitation, social coordination, and risky choice. Beyond validation, we uncover novel interactions among optimism, reward discretization, and temporal discounting. Specifically, we identify conditions under which choice hysteresis and path-dependent strategies emerge, suggesting that perseveration results from neural reward discretization rather than constituting an independent heuristic. We further reveal "individual dilemmas": circumstances where agents gravitate toward suboptimal yet stable strategies, offering a mechanistic explanation for incoherent choice patterns. Our framework bridges neuroscience, psychology, and collective behavior, enabling empirically testable hypotheses about how cognitive biases propagate from individual cognition to social outcomes in complex environments.

neuroscience↗

Moderate confirmation bias enhances collective decision-making

Humans tend to give more weight to information confirming their beliefs than to information that disconfirms them. Nevertheless, this apparent irrationality has been shown to improve individual decision-making under uncertainty. However, little is known about this bias impact on collective decision-making. Here, we investigate the conditions under which confirmation bias is beneficial or detrimental to collective decision-making. To do so, we develop a Collective Asymmetric Reinforcement Learning (CARL) model in which artificial agents observe others actions and rewards, and update this information asymmetrically. We use agent-based simulations to study how confirmation bias affects collective performance on a two-armed bandit task, and how resource scarcity, group size and bias strength modulate this effect. We find that a confirmation bias benefits group learning across a wide range of resource-scarcity conditions. Moreover, we discover that, past a critical bias strength, resource abundance favors the emergence of two different performance regimes, one of which is suboptimal. In addition, we find that this regime bifurcation comes with polarization in small groups of agents. Overall, our results suggest the existence of an optimal, moderate level of confirmation bias for collective decision-making. AUTHOR SUMMARYWhen we give more weight to information that confirms our existing beliefs, it typically has a negative impact on learning and decision-making. However, our study shows that moderate confirmation bias can actually improve collective decision-making when multiple reinforcement learning agents learn together in a social context. This finding has important implications for policymakers who engage in fighting against societal polarization and the spreading of misinformation. It can also inspire the development of artificial, distributed learning algorithms. Based on our research, we recommend not directly targeting confirmation bias but instead focusing on its underlying factors, such as group size, individual incentives, and the interactions between bias and the environment (such as filter bubbles).

animal behavior and cognition↗