EconPapers    
Economics at your fingertips  
 

Is Deep Hedging Reinforcement Learning?

Fr\'ed\'eric Godin

Papers from arXiv.org

Abstract: The deep hedging framework of Buehler et al. (2019) trains a neural network policy, via Monte Carlo simulation of price paths and stochastic gradient descent, to minimize a risk measure applied to the terminal hedging error. In a recent stream of papers, my coauthors and I have described this technique as reinforcement learning (RL). Several peers have, on occasion, expressed the view that deep hedging does not constitute genuine RL, on two grounds, among others: first, that because feedback is generated only at the terminal date, with no intermediate reward signal, the method cannot constitute genuine RL; and second, that the absence of a value function, a Bellman equation, temporal-difference (TD) learning, and an explicit exploration mechanism disqualifies the method from the RL category altogether, so that it should instead be labeled a neural-network method for stochastic optimal control. The present note argues instead that both objections rest on an unduly narrow, TD-centric reading of what constitutes RL, and that once RL is understood, as it is in the standard references of the field, to include Monte Carlo policy-gradient methods and direct (actor-only) policy search as first-class members, the deep hedging algorithm of Buehler et al. (2019) falls squarely within the RL umbrella.

Date: 2026-07, Revised 2026-07
References: Add references at CitEc
Citations:

Downloads: (external link)
https://arxiv.org/pdf/2607.13353 Latest version (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:arx:papers:2607.13353

Access Statistics for this paper

More papers in Papers from arXiv.org
Bibliographic data for series maintained by arXiv administrators ().

 
Page updated 2026-07-20
Handle: RePEc:arx:papers:2607.13353