Is Deep Hedging Reinforcement Learning?
Fr\'ed\'eric Godin
Papers from arXiv.org
Abstract:
The deep hedging framework of Buehler et al. (2019) trains a neural network policy, via Monte Carlo simulation of price paths and stochastic gradient descent, to minimize a risk measure applied to the terminal hedging error. In a recent stream of papers, my coauthors and I have described this technique as reinforcement learning (RL). Several peers have, on occasion, expressed the view that deep hedging does not constitute genuine RL, on two grounds, among others: first, that because feedback is generated only at the terminal date, with no intermediate reward signal, the method cannot constitute genuine RL; and second, that the absence of a value function, a Bellman equation, temporal-difference (TD) learning, and an explicit exploration mechanism disqualifies the method from the RL category altogether, so that it should instead be labeled a neural-network method for stochastic optimal control. The present note argues instead that both objections rest on an unduly narrow, TD-centric reading of what constitutes RL, and that once RL is understood, as it is in the standard references of the field, to include Monte Carlo policy-gradient methods and direct (actor-only) policy search as first-class members, the deep hedging algorithm of Buehler et al. (2019) falls squarely within the RL umbrella.
Date: 2026-07, Revised 2026-07
References: Add references at CitEc
Citations:
Downloads: (external link)
https://arxiv.org/pdf/2607.13353 Latest version (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:arx:papers:2607.13353
Access Statistics for this paper
More papers in Papers from arXiv.org
Bibliographic data for series maintained by arXiv administrators ().