Skip to main navigation Skip to search Skip to main content

Curriculum-Enhanced Reinforcement Learning for Robust Humanoid Locomotion

  • Harbin Institute of Technology
  • Eastern Institute of Technology, Ningbo

Research output: Contribution to journalArticlepeer-review

Abstract

The control of humanoid locomotion remains one of the most formidable challenges in robotics. Conventional model-based approaches not only rely extensively on manual design but also exhibit limited generalization across diverse tasks and environments. To overcome these limitations, a curriculum-enhanced reinforcement learning framework is proposed in this work for training robust locomotion. Given the strong temporal dependencies inherent in humanoid locomotion, Mamba-2 is employed as the backbone of the Actor-Critic network, enabling efficient temporal modeling of historical observations and content-based reasoning with linear spatiotemporal complexity. Furthermore, to eliminate the tendency of humanoid robots to favor low-speed motion in order to maintain gait stability and balance, which often leads to inaccurate velocity command tracking, a command-guided curriculum learning (CGCL) method is proposed to improve responsiveness to velocity commands. Experimental results demonstrate that the proposed Mamba-based framework achieves state-of-the-art (SOTA) performance, while CGCL significantly enhances the accuracy of velocity command tracking. Moreover, sim-to-real transfer experiments confirm both the robustness and the seamless deployability of the proposed framework on physical humanoid platforms. Note to Practitioners - Humanoid robots offer superior adaptability and performance in human-centered, unstructured environments, but controlling their whole-body locomotion remains challenging. Most existing methods rely on model-based designs that require extensive manual tuning, limiting their flexibility. This work proposes a novel curriculum-enhanced reinforcement learning framework that improves robustness and generalization for humanoid locomotion. On one hand, the framework utilizes Mamba-2 as the backbone of the Actor-Critic architecture, improving computational efficiency compared with current Transformer-based methods while retaining the ability to reason over complex movement sequences. On the other hand, a command-guided curriculum learning (CGCL) method is proposed to enhance the ability of the robot to accurately track velocity commands. Experimental results demonstrate that the proposed method exhibits strong robustness and precise command-tracking capabilities, highlighting its significant potential for real-world applications in humanoid locomotion.

Original languageEnglish
Pages (from-to)5779-5789
Number of pages11
JournalIEEE Transactions on Automation Science and Engineering
Volume23
DOIs
StatePublished - 2026

Keywords

  • Reinforcement learning
  • curriculum learning
  • end-to-end control
  • humanoid locomotion
  • learning-based control

Fingerprint

Dive into the research topics of 'Curriculum-Enhanced Reinforcement Learning for Robust Humanoid Locomotion'. Together they form a unique fingerprint.

Cite this