Abstract
Maneuverable threat poses significant risks to the operational safety of spacecraft in orbit, thereby jeopardizing mission success in the contested space environment. To address the coordinated orbital game between multiple pursuers (spacecraft) and one evader (threat), this article proposes a reward-centering multiagent deep deterministic policy gradient (MADDPG) algorithm, which has twofold. First, a relative dynamical model considering J2 perturbation is established to characterize orbital interactions among multiple players, laying a physics-based foundation for decision-making. Second, an improved multiagent reinforcement learning (RL) framework is constructed with two components. For individual spacecraft of multiple pursuers, a state-driven exploration-exploitation tradeoff mechanism is incorporated to avoid suboptimal strategy convergence. For the spacecraft group (namely, multiple pursuers), reward functions are designed to align individual behaviors with collective objectives. Furthermore, a reward-centering mechanism along with theoretical analysis is introduced to dynamically calibrate the reward baseline, which stabilizes policy update via reducing gradient variance. Simulation results show that the proposed algorithm enables spacecraft to autonomously develop coordinated strategies, ultimately confining the threat within a predefined interception threshold.
| Original language | English |
|---|---|
| Pages (from-to) | 10929-10938 |
| Number of pages | 10 |
| Journal | IEEE Internet of Things Journal |
| Volume | 13 |
| Issue number | 6 |
| DOIs | |
| State | Published - 15 Mar 2026 |
Keywords
- Game-based control
- multiagent system
- reinforcement learning (RL)
- spacecraft control
Fingerprint
Dive into the research topics of 'Reward-Centering Deep Reinforcement Learning for Multiagent Coordinated Interception in Orbital Pursuit-Evasion Game'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver