TY - GEN
T1 - Intelligent Reentry Guidance with Dynamic No-Fly Zones Based on Deep Reinforcement Learning
AU - Jiang, Qingji
AU - Wang, Xiaogang
AU - Li, Yu
N1 - Publisher Copyright:
© 2024, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2024
Y1 - 2024
N2 - Aimed at avoiding multiple dynamic no-fly zones and satisfying path constraints and terminal constraints in the reentry process of hypersonic glide vehicles, intelligent reentry guidance based on deep reinforcement learning is developed. Firstly, the guidance is decoupled as longitudinal guidance and lateral guidance. The lateral guidance provides the sign of the bank angle to adjust the heading direction while the longitudinal guidance outputs the magnitude of the bank angle through the artificial intelligence interface. Then, the reentry guidance simulation is mapped to a Markov Decision Process, in which the essential elements including state, action, and reward are defined or designed adaptively. Finally, the policy neural network is trained by the twin delayed deep deterministic policy gradient (TD3) algorithm. By selecting proper hyperparameters and network architecture, the policy neural network is able to converge. Simulations imply that under the influence of dynamic no-fly zones, initial state errors, and kinds of online dispersion, the proposed guidance can avoid all the no-fly zones and reach the target accurately with all the satisfied path constraints.
AB - Aimed at avoiding multiple dynamic no-fly zones and satisfying path constraints and terminal constraints in the reentry process of hypersonic glide vehicles, intelligent reentry guidance based on deep reinforcement learning is developed. Firstly, the guidance is decoupled as longitudinal guidance and lateral guidance. The lateral guidance provides the sign of the bank angle to adjust the heading direction while the longitudinal guidance outputs the magnitude of the bank angle through the artificial intelligence interface. Then, the reentry guidance simulation is mapped to a Markov Decision Process, in which the essential elements including state, action, and reward are defined or designed adaptively. Finally, the policy neural network is trained by the twin delayed deep deterministic policy gradient (TD3) algorithm. By selecting proper hyperparameters and network architecture, the policy neural network is able to converge. Simulations imply that under the influence of dynamic no-fly zones, initial state errors, and kinds of online dispersion, the proposed guidance can avoid all the no-fly zones and reach the target accurately with all the satisfied path constraints.
KW - Artificial intelligence
KW - Deep reinforcement learning
KW - Hypersonic glide vehicle
KW - No-fly zones
KW - Reentry guidance
UR - https://www.scopus.com/pages/publications/85180632994
U2 - 10.1007/978-3-031-42515-8_20
DO - 10.1007/978-3-031-42515-8_20
M3 - 会议稿件
AN - SCOPUS:85180632994
SN - 9783031425141
T3 - Mechanisms and Machine Science
SP - 291
EP - 313
BT - Computational and Experimental Simulations in Engineering - Proceedings of ICCES 2023—Volume 1
A2 - Li, Shaofan
PB - Springer Science and Business Media B.V.
T2 - 29th International Conference on Computational and Experimental Engineering and Sciences, ICCES 2023
Y2 - 26 May 2023 through 29 May 2023
ER -