Skip to main navigation Skip to search Skip to main content

Autonomous Defensive Decision-Making in Orbital Target–Attacker–Defender Engagements via Proximal Policy Optimization

  • Xiaoran Li
  • , Xu Tang*
  • , Dong Ye
  • , Jiahao Yang
  • , Yibin He
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Compared to conventional pursuit-evasion scenarios, Target–Attacker–Defender (TAD) orbital games exhibit higher complexity due to coupled multi-body kinematics. To address autonomous defense in these TAD orbital games, this study formulates an intelligent decision-making framework driven by Proximal Policy Optimization (PPO). The established TAD game model incorporates practical operational constraints, primarily encompassing bounded impulsive thrusts, discrete decision epochs, and finite engagement horizons. Within this highly dynamic environment, the Attacker executes a multi-impulse interception strategy optimized via Sequential Quadratic Programming (SQP), serving as a stringent adversarial baseline. Against such optimized adversarial maneuvers, the Defender agent is trained via PPO. To overcome exploration bottlenecks inherent in sparse-reward environments, the training is guided by a customized composite reward function that explicitly captures both relative kinematics and collinear spatial configurations to steer policy learning. Simulation results demonstrate that the trained Defender autonomously synthesizes and executes non-coplanar maneuvers that promote collinear interposition. This emergent tactical intelligence effectively obstructs the adversarial approach vector, systematically compelling the SQP-driven Attacker to abort its offensive trajectory, thereby improving Target survivability throughout the engagement.

Original languageEnglish
JournalAdvances in Astronautics
DOIs
StateAccepted/In press - 2026

Keywords

  • Orbital game
  • Proximal policy optimization
  • Reinforcement learning
  • Target defense

Fingerprint

Dive into the research topics of 'Autonomous Defensive Decision-Making in Orbital Target–Attacker–Defender Engagements via Proximal Policy Optimization'. Together they form a unique fingerprint.

Cite this