Robust policy iteration for the continuous-time stochastic H∞ control problem with unknown dynamics
Zhongshi Sun and
Guangyan Jia
Mathematics and Computers in Simulation (MATCOM), 2026, vol. 241, issue PA, 430-448
Abstract:
In this article, we study a continuous-time stochastic H∞ control problem using reinforcement learning (RL) techniques, which can be formulated as solving a stochastic linear-quadratic two-person zero-sum differential game (LQZSG). First, we propose a PI-based RL algorithm that iteratively solves the stochastic game algebraic Riccati equation using collected state and control data, with all system dynamic information unknown. Notably, the algorithm requires data collection only once during the iteration process. We then provide a convergence proof of the RL algorithm and analyze the robustness of the inner and outer loops of the PI algorithm, demonstrating that when the iteration error is within a certain range, the algorithm converges to a small neighborhood of the saddle point of the stochastic LQZSG problem. Finally, we validate the effectiveness of the proposed RL algorithm through two simulation examples.
Keywords: Reinforcement learning; Policy iteration; Stochastic H∞ control; Linear-quadratic game; Robustness (search for similar items in EconPapers)
Date: 2026
References: View references in EconPapers View complete reference list from CitEc
Citations:
Downloads: (external link)
http://www.sciencedirect.com/science/article/pii/S0378475425003817
Full text for ScienceDirect subscribers only
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:eee:matcom:v:241:y:2026:i:pa:p:430-448
DOI: 10.1016/j.matcom.2025.09.009
Access Statistics for this article
Mathematics and Computers in Simulation (MATCOM) is currently edited by Robert Beauwens
More articles in Mathematics and Computers in Simulation (MATCOM) from Elsevier
Bibliographic data for series maintained by Catherine Liu ().