Deterministic dynamics of distributional multi-agent reinforcement learning
Clémence Bergerot,
Pawel Romanczuk and
Wolfram Barfuss
PLOS Computational Biology, 2026, vol. 22, issue 9, 1-24
Abstract:
Understanding how cognition shapes behavior across contexts remains a fundamental challenge for many disciplines. In particular, for the optimism heuristic–i.e., the tendency to overweight positive (relative to negative) information–knowledge remains fragmented, with models developed in specific domains in isolation. Here, we present a unifying computational framework by deriving the deterministic dynamics of distributional multi-agent reinforcement learning. Our approach discretizes return distributions through a finite set of neurons, consistent with recent empirical findings on distributional coding in the brain. We validate our framework by reproducing established results across three iconic domains spanning individual bandit choice under resource variability, social coordination, and risky choice. Beyond validation, we uncover novel interactions among optimism, return discretization, and temporal discounting. Specifically, we identify conditions under which return discretization generates choice hysteresis and, in extreme parameter regimes, inescapable perseveration. We further reveal “individual dilemmas”: circumstances where agents gravitate toward suboptimal yet stable strategies, offering a mechanistic explanation for incoherent choice patterns. Our framework bridges neuroscience, psychology, and collective behavior, enabling empirically testable hypotheses about how cognitive biases propagate from individual cognition to social outcomes in complex environments.Author summary: How do cognitive heuristics shape individual decisions and the collective outcomes that emerge when many agents interact? This question connects psychology, neuroscience, and artificial intelligence, yet computational models have developed separately across these fields. Here we present DDRL, a framework that combines distributional reinforcement learning, in which agents learn the full spread of possible outcomes rather than their average, with deterministic learning dynamics that yield mathematically tractable trajectories in strategy space. DDRL is grounded in neuroscientific evidence: dopaminergic neurons encode outcome distributions rather than expected values alone. We illustrate DDRL with an optimism heuristic across three domains: bandit choice under resource variability, social coordination, and intertemporal risky choice. In each case, DDRL reproduces established results. Beyond replication, DDRL makes two novel predictions. Return discretization generates bistable regimes, in which agents settle into qualitatively different behaviors depending on their history. Under certain combinations of optimism and temporal discounting, agents can become locked into an individual dilemma, in which the better strategy is no longer a stable attractor. Together, these results provide a mechanistic route from neural reward encoding to emergent collective behavior, connecting individual cognitive heuristics to social outcomes in a biologically grounded and analytically tractable framework.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014723 (text/html)
https://journals.plos.org/ploscompbiol/article/fil ... 14723&type=printable (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:plo:pcbi00:1014723
DOI: 10.1371/journal.pcbi.1014723
Access Statistics for this article
More articles in PLOS Computational Biology from Public Library of Science
Bibliographic data for series maintained by ploscompbiol ().