Abstract
Compared to conventional pursuit-evasion scenarios, Target–Attacker–Defender (TAD) orbital games exhibit higher complexity due to coupled multi-body kinematics. To address autonomous defense in these TAD orbital games, this study formulates an intelligent decision-making framework driven by Proximal Policy Optimization (PPO). The established TAD game model incorporates practical operational constraints, primarily encompassing bounded impulsive thrusts, discrete decision epochs, and finite engagement horizons. Within this highly dynamic environment, the Attacker executes a multi-impulse interception strategy optimized via Sequential Quadratic Programming (SQP), serving as a stringent adversarial baseline. Against such optimized adversarial maneuvers, the Defender agent is trained via PPO. To overcome exploration bottlenecks inherent in sparse-reward environments, the training is guided by a customized composite reward function that explicitly captures both relative kinematics and collinear spatial configurations to steer policy learning. Simulation results demonstrate that the trained Defender autonomously synthesizes and executes non-coplanar maneuvers that promote collinear interposition. This emergent tactical intelligence effectively obstructs the adversarial approach vector, systematically compelling the SQP-driven Attacker to abort its offensive trajectory, thereby improving Target survivability throughout the engagement.
| Original language | English |
|---|---|
| Journal | Advances in Astronautics |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Orbital game
- Proximal policy optimization
- Reinforcement learning
- Target defense
Fingerprint
Dive into the research topics of 'Autonomous Defensive Decision-Making in Orbital Target–Attacker–Defender Engagements via Proximal Policy Optimization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver