Abstract
Traditional ascent trajectory optimization methods for hypersonic vehicles suffer from a high computational burden. This paper proposes an intelligent optimization algorithm based on deep reinforcement learning, where the offline-trained neural network enables fast trajectory generation and online deployment while maintaining all terminal, control, and path constraints. First, the trajectory optimization model—including ascent dynamics, program angle form, and multiple constraints—is established. Then, the trajectory optimization problem is formulated as a Markov decision process, rendering it solvable by reinforcement learning. To accelerate training, the twin delayed deep deterministic policy gradient (TD3) algorithm is improved via behavioral cloning and learning from demonstrations (fD). Specifically, an expert policy is pretrained via behavioral cloning from the optimization dataset, which can warm-start the TD3′s Actor. As a result, the early-stage exploration for successful samples is effectively enhanced. Additionally, a supervised auxiliary loss is incorporated into the Actor′s update to strengthen the reward signal and improve TD3 convergence. Finally, simulations demonstrate that the proposed behavioral cloning-augmented TD3 from demonstrations (BC-TD3fD) algorithm converges rapidly, whereas the original TD3 fails. Compared with standard baselines, BC-TD3fD reduces the average offline optimization time from 19.268 to 0.149 s, achieving a substantial acceleration over sequential quadratic programming, and decreases the online per-step computation time to 3.204 ms, representing a 94.2% reduction relative to model predictive control based on quadratic programming, thereby ensuring deterministic real-time performance.
| Original language | English |
|---|---|
| Article number | 7263139 |
| Journal | International Journal of Aerospace Engineering |
| Volume | 2026 |
| Issue number | 1 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- ascent trajectory optimization
- behavioral cloning
- hypersonic vehicles
- learning from demonstrations
- reinforcement learning
Fingerprint
Dive into the research topics of 'Intelligent Ascent Trajectory Optimization via Behavioral Cloning-Augmented Reinforcement Learning From Demonstrations'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver