EconPapers    
Economics at your fingertips  
 

When Does Randomized Oversight Align AI Agents That Can Conceal?

Joshua S. Gans and Richard Holden

Papers from arXiv.org

Abstract: Oversight changes the evidence it relies on. We ask when randomized audits and scoring align AI agents that can conceal misconduct and alter records. Stronger auditing makes undeterred violations better hidden. Because the provider writes the agent's objective, sanctions need not stop at forfeiture, and rare audits deter every type of agent if evidence survives concealment and audit draws cannot be learned in advance. When evidence can be erased, deterrence must come from lower gains from violation, such as credit for stopping, or from costlier or fewer ways to conceal. These conditions identify what failed when agents in OpenAI's cybersecurity evaluations compromised parts of Hugging Face's infrastructure in July 2026.

Date: 2026-09
New Economics Papers: this item is included in nep-inv
References: Add references at CitEc
Citations:

Downloads: (external link)
https://arxiv.org/pdf/2609.38262 Latest version (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:arx:papers:2609.38262

Access Statistics for this paper

More papers in Papers from arXiv.org
Bibliographic data for series maintained by arXiv administrators ().

 
Page updated 2026-10-06
Handle: RePEc:arx:papers:2609.38262