Delayed reward information is underweighted in reinforcement learning with dispersed feedback
Miruna Cotet,
David Poensgen and
Ian Krajbich
PLOS Computational Biology, 2026, vol. 22, issue 6, 1-25
Abstract:
Learning is fundamental to adaptive behavior. In the typical learning task, each action is associated with only one outcome, which could be immediate or delayed. However, actions often have multiple consequences that unfold over time. Here, we used behavioral and eye-tracking experiments to study how people learn when their choices yield both immediate and delayed reward information. Importantly, the rewards themselves were all delivered at the end of the study so there was no reason to weight immediate and delayed reward information differently. Instead, we found that our subjects overweighted immediate reward information. Moreover, this bias increased over the course of the experiment and was still present when learning from others’ choices. The gaze data reveal mixed evidence that subjects looked more at immediate vs. delayed feedback, and across subjects, the relative dwell proportion did not predict the behavioral bias. Our results indicate that people prioritize not just immediate rewards, but immediate reward information. Unlike temporal discounting, this form of impatience is a clear mistake and leads to objectively worse outcomes.Author summary: Decisions often produce several pieces of feedback. Some occur right away, others occur later. Although these signals can be equally useful, people may not treat them that way. In our study, we asked whether humans learn differently from information that arrives immediately versus information that arrives after a delay, even when both are equally valuable. Using a combination of behavioral experiments and eye-tracking, we found that people consistently placed too much emphasis on the immediate feedback. This tendency grew stronger over time and even appeared when participants were learning by watching others’ choices. Our findings show that people aren’t just drawn to immediate rewards—they are also overly affected by immediate information about those rewards. Unlike classic impatience, where immediate rewards are often genuinely better, immediate information provides no benefit in our setting and overweighting it leads to poor decisions. Our study reveals a systematic and costly quirk in how people learn, and offers an explanation for why people often seem so shortsighted.
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014459 (text/html)
https://journals.plos.org/ploscompbiol/article/fil ... 14459&type=printable (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:plo:pcbi00:1014459
DOI: 10.1371/journal.pcbi.1014459
Access Statistics for this article
More articles in PLOS Computational Biology from Public Library of Science
Bibliographic data for series maintained by ploscompbiol ().