TY - GEN
T1 - APF-Guided PPO for Obstacle Avoidance and Trajectory Planning of Manipulators
AU - Yao, Bowei
AU - Liu, Zhuang
AU - Zhu, Fuxing
AU - Wu, Yuqiang
AU - Tuo, Liheng
AU - Liu, Jianxing
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Trajectory planning for robotic manipulators in complex, dynamic environments presents a fundamental challenge. While Deep Reinforcement Learning (DRL) offers a promising solution for real-time decision-making, standard algorithms such as Proximal Policy Optimization (PPO) often struggle with sparse rewards and inefficient exploration in high-dimensional continuous action spaces. To address these limitations, this work proposes a novel Dynamic Artificial Potential Field-Guided PPO (DAPF-PPO). We introduce a Dynamic Potential-Based Reward Shaping (DPBRS) mechanism that integrates prior physical knowledge from Artificial Potential Fields (APF) into the RL reward structure. During the early training stages, dense APF-based rewards effectively guide the agent, substantially accelerating the initial learning phase. As training progresses, the guidance weight gradually attenuates, allowing the PPO agent to dominate decision-making and explore globally optimal policies without bias. Simulations utilizing a Franka Emika Panda manipulator within the MuJoCo physics engine demonstrate the effectiveness of the proposed framework. Compared with standard PPO, DAPF-PPO significantly suppresses policy oscillations during early exploration, mitigates the sparse reward problem, and achieves higher success rates as well as smoother trajectories in dynamic obstacle avoidance tasks.
AB - Trajectory planning for robotic manipulators in complex, dynamic environments presents a fundamental challenge. While Deep Reinforcement Learning (DRL) offers a promising solution for real-time decision-making, standard algorithms such as Proximal Policy Optimization (PPO) often struggle with sparse rewards and inefficient exploration in high-dimensional continuous action spaces. To address these limitations, this work proposes a novel Dynamic Artificial Potential Field-Guided PPO (DAPF-PPO). We introduce a Dynamic Potential-Based Reward Shaping (DPBRS) mechanism that integrates prior physical knowledge from Artificial Potential Fields (APF) into the RL reward structure. During the early training stages, dense APF-based rewards effectively guide the agent, substantially accelerating the initial learning phase. As training progresses, the guidance weight gradually attenuates, allowing the PPO agent to dominate decision-making and explore globally optimal policies without bias. Simulations utilizing a Franka Emika Panda manipulator within the MuJoCo physics engine demonstrate the effectiveness of the proposed framework. Compared with standard PPO, DAPF-PPO significantly suppresses policy oscillations during early exploration, mitigates the sparse reward problem, and achieves higher success rates as well as smoother trajectories in dynamic obstacle avoidance tasks.
KW - Artificial Potential Field
KW - Deep Reinforcement Learning
KW - Obstacle Avoidance
KW - Proximal Policy Optimization
KW - Trajectory Planning
UR - https://www.scopus.com/pages/publications/105043889142
U2 - 10.1109/CCDC69976.2026.11560488
DO - 10.1109/CCDC69976.2026.11560488
M3 - 会议稿件
AN - SCOPUS:105043889142
T3 - 38th Chinese Control and Decision Conference, CCDC 2026
SP - 5368
EP - 5373
BT - 38th Chinese Control and Decision Conference, CCDC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 38th Chinese Control and Decision Conference, CCDC 2026
Y2 - 15 May 2026 through 18 May 2026
ER -