Whistleblowers can contain the unethical externalities of human-AI delegation
Zoe Purcell,
Nils Köbis,
Andrew Samuel and
Jean-François Bonnefon
Additional contact information
Zoe Purcell: LaPsyDÉ - UMR 8240 - Laboratoire de psychologie du développement et de l'éducation de l'enfant - CNRS - Centre National de la Recherche Scientifique - UPCité - Université Paris Cité
Nils Köbis: Universität Duisburg-Essen = University of Duisburg-Essen [Essen]
Andrew Samuel: Loyola University [Maryland, Baltimore]
Jean-François Bonnefon: TSE-R - Toulouse School of Economics - UT Capitole - Université Toulouse Capitole - Comue de Toulouse - Communauté d'universités et établissements de Toulouse - EHESS - École des hautes études en sciences sociales - CNRS - Centre National de la Recherche Scientifique - INRAE - Institut National de Recherche pour l’Agriculture, l’Alimentation et l’Environnement, TSM - Toulouse School of Management Research - UT Capitole - Université Toulouse Capitole - Comue de Toulouse - Communauté d'universités et établissements de Toulouse - CNRS - Centre National de la Recherche Scientifique - TSM - Toulouse School of Management - UT Capitole - Université Toulouse Capitole - Comue de Toulouse - Communauté d'universités et établissements de Toulouse, CNRS - Centre National de la Recherche Scientifique
Post-Print from HAL
Abstract:
Prior work using controlled principal-agent experiments suggests two risks of delegating tasks to AI systems: human principals are more likely to request profit-maximizing misconduct from AI agents than from human agents, and AI agents are more likely to comply. Here we test whether third-party observers can contain the resulting harm. In an incentivized die-reporting paradigm, principals instructed either a human or an AI agent how strongly to prioritize profit over accuracy, creating potential financial harm to a charity. We first confirm, with human principals (N = 600) and three large language models as AI agents, that delegation to AI produces larger negative externalities than delegation to humans. We then study observers who could pay a personal cost to flag a principal's instruction, cancelling the principal's gain in favor of the charity, as a laboratory analogue of whistleblowing. In this observer study (N = 300), the probability of flagging increased with how unethical the principal's request was, but did not depend on whether the request was directed to a human or an AI agent. Because principals made more unethical requests under AI delegation, flagging was more frequent under AI delegation. When combined with agent behavior, this increase in flagging fully neutralized the negative externalities of AI delegation in our experimental setting. These findings support institutional protections for whistleblowers as one potential organizational safeguard against the harms of human-AI delegation.
Keywords: Game theory; Whistleblowing; Delegation; Human-AI interaction (search for similar items in EconPapers)
Date: 2026-07-16
References: Add references at CitEc
Citations:
Published in Proceedings of the National Academy of Sciences of the United States of America, 2026, 123 (29), pp.e2536668123. ⟨10.1073/pnas.2536668123⟩
There are no downloads for this item, see the EconPapers FAQ for hints about obtaining it.
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:hal:journl:hal-05745755
DOI: 10.1073/pnas.2536668123
Access Statistics for this paper
More papers in Post-Print from HAL
Bibliographic data for series maintained by CCSD ().