TY - GEN
T1 - BarlowWalk
T2 - 24th IEEE-RAS International Conference on Humanoid Robots, Humanoids 2025
AU - Huang, Haodong
AU - Sun, Shilong
AU - Wang, Yuanpeng
AU - Li, Chiyao
AU - Huang, Hailin
AU - Xu, Wenfu
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Reinforcement learning (RL), driven by datadriven methods, has become an effective solution for robot leg motion control problems. However, the mainstream RL methods for bipedal robot terrain traversal, such as teacher-student policy knowledge distillation, suffer from long training times, which limit development efficiency. To address this issue, this paper proposes BarlowWalk, an improved Proximal Policy Optimization (PPO) method integrated with selfsupervised representation learning. This method employs the Barlow Twins algorithm to construct a decoupled latent space, mapping historical observation sequences into low-dimensional representations and implementing self-supervision. Meanwhile, the actor requires only proprioceptive information to achieve self-supervised learning over continuous time steps, significantly reducing the dependence on external terrain perception. Simulation experiments demonstrate that this method has significant advantages in complex terrain scenarios. To enhance the credibility of the evaluation, this study compares BarlowWalk with advanced algorithms through comparative tests, and the experimental results verify the effectiveness of the proposed method.
AB - Reinforcement learning (RL), driven by datadriven methods, has become an effective solution for robot leg motion control problems. However, the mainstream RL methods for bipedal robot terrain traversal, such as teacher-student policy knowledge distillation, suffer from long training times, which limit development efficiency. To address this issue, this paper proposes BarlowWalk, an improved Proximal Policy Optimization (PPO) method integrated with selfsupervised representation learning. This method employs the Barlow Twins algorithm to construct a decoupled latent space, mapping historical observation sequences into low-dimensional representations and implementing self-supervision. Meanwhile, the actor requires only proprioceptive information to achieve self-supervised learning over continuous time steps, significantly reducing the dependence on external terrain perception. Simulation experiments demonstrate that this method has significant advantages in complex terrain scenarios. To enhance the credibility of the evaluation, this study compares BarlowWalk with advanced algorithms through comparative tests, and the experimental results verify the effectiveness of the proposed method.
UR - https://www.scopus.com/pages/publications/105022186590
U2 - 10.1109/Humanoids65713.2025.11203059
DO - 10.1109/Humanoids65713.2025.11203059
M3 - 会议稿件
AN - SCOPUS:105022186590
T3 - IEEE-RAS International Conference on Humanoid Robots
SP - 906
EP - 913
BT - 2025 IEEE-RAS 24th International Conference on Humanoid Robots, Humanoids 2025
PB - IEEE Computer Society
Y2 - 30 September 2025 through 2 October 2025
ER -