TY - GEN
T1 - Pursuit-Evasion Games for Multiple Unmanned Surface Vehicles Based on LD-MADDPG
AU - Ma, Yuange
AU - Zi, Lin
AU - Yu, Qianbing
AU - Zhuang, Yufei
AU - Jin, Liangjun
AU - Li, Lang
AU - Huang, Haibin
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper proposes a cooperative strategy based on the LD-MADDPG(Long Short-Term Memory-Differential Latency Gating Mechanism- Multi-Agent Deep Deterministic Policy Gradient) framework for the multi unmanned surface vessel(USV) pursuit game task. To address the problem of insufficient utilization of temporal information in individual decision-making and sparse group collaboration, Long Short-Term Memory(LSTM) is embedded in the network, enabling each USV to implicitly maintain a dynamic contextual memory, thereby making more continuous and robust maneuvering decisions in partially observable water surface environments. Further, to avoid the curse of dimensionality caused by the expansion of memory vectors over time, this paper introduces a differential saliency gating mechanism after the LSTM output: this module assigns importance weights to the implicit states at different historical moments, suppresses redundant information irrelevant to the current chase and escape phase, and retains only compact expressions sensitive to the critical situation of 'capture-escape'. This significantly enhances the efficiency of cluster collaboration. Simulation and measurement results show that the proposed algorithm has better convergence speed and average reward than the original Multi-Agent Deep Deterministic Policy Gradient(MADDPG).
AB - This paper proposes a cooperative strategy based on the LD-MADDPG(Long Short-Term Memory-Differential Latency Gating Mechanism- Multi-Agent Deep Deterministic Policy Gradient) framework for the multi unmanned surface vessel(USV) pursuit game task. To address the problem of insufficient utilization of temporal information in individual decision-making and sparse group collaboration, Long Short-Term Memory(LSTM) is embedded in the network, enabling each USV to implicitly maintain a dynamic contextual memory, thereby making more continuous and robust maneuvering decisions in partially observable water surface environments. Further, to avoid the curse of dimensionality caused by the expansion of memory vectors over time, this paper introduces a differential saliency gating mechanism after the LSTM output: this module assigns importance weights to the implicit states at different historical moments, suppresses redundant information irrelevant to the current chase and escape phase, and retains only compact expressions sensitive to the critical situation of 'capture-escape'. This significantly enhances the efficiency of cluster collaboration. Simulation and measurement results show that the proposed algorithm has better convergence speed and average reward than the original Multi-Agent Deep Deterministic Policy Gradient(MADDPG).
KW - LSTM
KW - MADDPG
KW - Pursuit-Evasion Game
KW - USV
UR - https://www.scopus.com/pages/publications/105043930157
U2 - 10.1109/CCDC69976.2026.11560703
DO - 10.1109/CCDC69976.2026.11560703
M3 - 会议稿件
AN - SCOPUS:105043930157
T3 - 38th Chinese Control and Decision Conference, CCDC 2026
SP - 1711
EP - 1717
BT - 38th Chinese Control and Decision Conference, CCDC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 38th Chinese Control and Decision Conference, CCDC 2026
Y2 - 15 May 2026 through 18 May 2026
ER -