TY - GEN
T1 - Learning a faster locomotion gait for a quadruped robot with model-free deep reinforcement learning
AU - Hu, Biao
AU - Shao, Shibo
AU - Cao, Zhengcai
AU - Xiao, Qing
AU - Li, Qunzhi
AU - Ma, Chao
N1 - Publisher Copyright:
© 2019 IEEE.
PY - 2019/12
Y1 - 2019/12
N2 - Quadruped robots have great agility, flexibility and stability, which enables them to walk through uneven terrain. Motion control of legged robots is always a difficult problem. Previous approaches mostly either use a predefined gait that results in clumsy and unnatural behavior, or use reinforcement learning approach to generate a gait strategy that needs longtime computation and elegant network design. In this paper, we present an effective approach that uses deep reinforcement learning with prior knowledge to optimize the gait of quadruped robot. By using a specific quadruped robot walking gait as a priori knowledge, this approach adopts the technique of distributed proximal policy optimization to optimize the search for better gait. The proposed approach does not require modelling of complex robots, and has good network convergence speed and learning effect. Simulation results demonstrate that our proposed approach converges faster than other deep reinforcement learning methods without prior knowledge. Besides, our achieved gait has higher speed that is 50% faster than the trot gait without optimization.
AB - Quadruped robots have great agility, flexibility and stability, which enables them to walk through uneven terrain. Motion control of legged robots is always a difficult problem. Previous approaches mostly either use a predefined gait that results in clumsy and unnatural behavior, or use reinforcement learning approach to generate a gait strategy that needs longtime computation and elegant network design. In this paper, we present an effective approach that uses deep reinforcement learning with prior knowledge to optimize the gait of quadruped robot. By using a specific quadruped robot walking gait as a priori knowledge, this approach adopts the technique of distributed proximal policy optimization to optimize the search for better gait. The proposed approach does not require modelling of complex robots, and has good network convergence speed and learning effect. Simulation results demonstrate that our proposed approach converges faster than other deep reinforcement learning methods without prior knowledge. Besides, our achieved gait has higher speed that is 50% faster than the trot gait without optimization.
UR - https://www.scopus.com/pages/publications/85079071167
U2 - 10.1109/ROBIO49542.2019.8961651
DO - 10.1109/ROBIO49542.2019.8961651
M3 - 会议稿件
AN - SCOPUS:85079071167
T3 - IEEE International Conference on Robotics and Biomimetics, ROBIO 2019
SP - 1097
EP - 1102
BT - IEEE International Conference on Robotics and Biomimetics, ROBIO 2019
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2019 IEEE International Conference on Robotics and Biomimetics, ROBIO 2019
Y2 - 6 December 2019 through 8 December 2019
ER -