Skip to main navigation Skip to search Skip to main content

A Coverage-based multi-Agent reinforcement learning method for cooperative guidance

  • Xinran Zhang
  • , Fenghua He*
  • , Yu Yao
  • , Zhaochen Lin
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

The increasing speed and agility of modern aerial targets, coupled with the prevalence of sensor noise in real-world scenarios, pose significant challenges to traditional guidance methods. This highlights an urgent need for advanced and robust cooperative pursuit strategies for multi-vehicle systems. To address these challenges, this paper proposes a novel coverage-based multi-agent reinforcement learning guidance framework, designed to achieve highly efficient and robust cooperative guidance. Specifically, we first reformulate the pursuit task as a dynamic coverage optimization problem over the target’s predicted reachable region. We then introduce a Projection-Embedded Coverage Optimization RL (PECO-RL) framework, which enables efficient training within a simplified projected 2D environment while maintaining validation capability in complex 3D scenarios, significantly reducing training complexity and enhancing generalization. Furthermore, we present a Coverage-based Reinforcement Learning Guidance (CRL-G) method to promote coordinated behavior that is inherently more robust to observation noise, thereby achieving superior pursuit efficiency and robustness. The CRL-G method integrates a coverage-driven reward function to mitigate the inherent sparsity problem in pursuit tasks. An Actor-Critic network architecture with an adaptive feature extraction module employing learnable attention mechanisms is designed to dynamically accommodate varying numbers of flight vehicles. Extensive simulation experiments demonstrate that the proposed CRL-G method achieves superior performance compared to existing approaches in terms of guidance accuracy, success rate, robustness, and control effort. Furthermore, the proposed method exhibits significantly higher computational efficiency than optimization-based methods, highlighting its potential for real-time deployment.

Original languageEnglish
Article number111717
JournalAerospace Science and Technology
Volume173
DOIs
StatePublished - Jun 2026

Keywords

  • Attention mechanisms
  • Cooperative guidance
  • Coverage-Based optimization
  • High-maneuvering target
  • Multi-Agent reinforcement learning
  • Proximal policy optimization (PPO)

Fingerprint

Dive into the research topics of 'A Coverage-based multi-Agent reinforcement learning method for cooperative guidance'. Together they form a unique fingerprint.

Cite this