TY - GEN
T1 - DDPGRU
T2 - 4th Conference on Fully Actuated System Theory and Applications, FASTA 2025
AU - Zhou, Yi
AU - Guo, Chuanjun
AU - Zhang, Tianhao
AU - Li, Zijing
AU - Qiu, Jianbin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Deep reinforcement learning (DRL) empowers agents to learn complex behaviors through dynamic interaction with the environment, and has been widely applied in fields such as robotics, autonomous driving, and gaming. As one of the popular and widely applied DRL algorithms, Deep Deterministic Policy Gradient (DDPG) combines the powerful function approximation capabilities of neural networks with deterministic policy gradients, successfully addressing reinforcement learning challenges in continuous action spaces. However, DDPG often fails to capture the temporal dependencies of dynamic states in real-world environments, leading to suboptimal policies and poor generalization performance. In this paper, the Gated Recurrent Unit (GRU) is incorporated as the backbone network of the Actor to learn the temporal dependencies of state dynamics. Based on this enhancement, an improved algorithm called DDPGRU is proposed to address the limitation of insufficient temporal modeling capabilities. The key innovation of DDPGRU is that DDPG provides a robust framework for continuous control, while GRU enables the agent to simulate and leverage temporal patterns within the environment. Experimental results demonstrate that the proposed DDPGRU algorithm significantly outperforms the original DDPG baseline algorithm.
AB - Deep reinforcement learning (DRL) empowers agents to learn complex behaviors through dynamic interaction with the environment, and has been widely applied in fields such as robotics, autonomous driving, and gaming. As one of the popular and widely applied DRL algorithms, Deep Deterministic Policy Gradient (DDPG) combines the powerful function approximation capabilities of neural networks with deterministic policy gradients, successfully addressing reinforcement learning challenges in continuous action spaces. However, DDPG often fails to capture the temporal dependencies of dynamic states in real-world environments, leading to suboptimal policies and poor generalization performance. In this paper, the Gated Recurrent Unit (GRU) is incorporated as the backbone network of the Actor to learn the temporal dependencies of state dynamics. Based on this enhancement, an improved algorithm called DDPGRU is proposed to address the limitation of insufficient temporal modeling capabilities. The key innovation of DDPGRU is that DDPG provides a robust framework for continuous control, while GRU enables the agent to simulate and leverage temporal patterns within the environment. Experimental results demonstrate that the proposed DDPGRU algorithm significantly outperforms the original DDPG baseline algorithm.
KW - Actor-Critic
KW - Deep Deterministic Policy Gradient
KW - Deep Reinforcement Learning
KW - Gated Recurrent Unit
UR - https://www.scopus.com/pages/publications/105017609374
U2 - 10.1109/FASTA65681.2025.11138126
DO - 10.1109/FASTA65681.2025.11138126
M3 - 会议稿件
AN - SCOPUS:105017609374
T3 - Proceedings of the 4th Conference on Fully Actuated System Theory and Applications, FASTA 2025
SP - 338
EP - 342
BT - Proceedings of the 4th Conference on Fully Actuated System Theory and Applications, FASTA 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 4 July 2025 through 6 July 2025
ER -