Deep Reinforcement-Learning-Based Air-Combat-Maneuver Generation Framework

Mei, Junru; Li, Ge; Huang, Hesong

Deep Reinforcement-Learning-Based Air-Combat-Maneuver Generation Framework

Junru Mei, Ge Li and Hesong Huang ()
Additional contact information
Junru Mei: College of Systems Engineering, National University of Defense Technology, Changsha 410073, China
Ge Li: College of Systems Engineering, National University of Defense Technology, Changsha 410073, China
Hesong Huang: College of Systems Engineering, National University of Defense Technology, Changsha 410073, China

Mathematics, 2024, vol. 12, issue 19, 1-24

Abstract: With the development of unmanned aircraft and artificial intelligence technology, the future of air combat is moving towards unmanned and autonomous direction. In this paper, we introduce a new layered decision framework designed to address the six-degrees-of-freedom (6-DOF) aircraft within-visual-range (WVR) air-combat challenge. The decision-making process is divided into two layers, each of which is addressed separately using reinforcement learning (RL). The upper layer is the combat policy, which determines maneuvering instructions based on the current combat situation (such as altitude, speed, and attitude). The lower layer control policy then uses these commands to calculate the input signals from various parts of the aircraft (aileron, elevator, rudder, and throttle). Among them, the control policy is modeled as a Markov decision framework, and the combat policy is modeled as a partially observable Markov decision framework. We describe the two-layer training method in detail. For the control policy, we designed rewards based on expert knowledge to accurately and stably complete autonomous driving tasks. At the same time, for combat policy, we introduce a self-game-based course learning, allowing the agent to play against historical policies during training to improve performance. The experimental results show that the operational success rate of the proposed method against the game theory baseline reaches 85.7%. Efficiency was also outstanding, with an average 13.6% reduction in training time compared to the RL baseline.

Keywords: air combat; deep reinforcement learning; SAC; recurrent neural network (search for similar items in EconPapers)
JEL-codes: C (search for similar items in EconPapers)
Date: 2024
References: View complete reference list from CitEc
Citations:

Downloads: (external link)
https://www.mdpi.com/2227-7390/12/19/3020/pdf (application/pdf)
https://www.mdpi.com/2227-7390/12/19/3020/ (text/html)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:gam:jmathe:v:12:y:2024:i:19:p:3020-:d:1487505

Access Statistics for this article

Mathematics is currently edited by Ms. Emma He

More articles in Mathematics from MDPI
Bibliographic data for series maintained by MDPI Indexing Manager ().