Skip to main navigation Skip to search Skip to main content

Mastering table tennis with hierarchy: a reinforcement learning approach with progressive self-play training

  • Hongxu Ma
  • , Jianyin Fan
  • , Haoran Xu
  • , Qiang Wang*
  • *Corresponding author for this work
  • School of Astronautics, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Hierarchical Reinforcement Learning (HRL) is widely applied in various complex task scenarios. In complex tasks where simple model-free reinforcement learning struggles, hierarchical design allows for more efficient utilization of interactive data, significantly reducing training costs and improving training success rates. This study delves into the use of HRL based on the model-free policy layer to learn complex strategies for a robotic arm playing table tennis. Through processes such as pre-training, self-play training, and self-play training with top-level winning strategies, the robustness of the lower-level hitting strategies has been enhanced. Furthermore, a novel decay reward mechanism has been employed in the training of the higher-level agent to improve the win rate in adversarial matches against other methods. After pre-training and adversarial training, we achieved an average of 52 rally cycles for the forehand strategy and 48 rally cycles for the backhand strategy in testing. The high-level strategy training based on the decay reward mechanism resulted in an advantageous score when competing against other strategies.

Original languageEnglish
Article number562
JournalApplied Intelligence
Volume55
Issue number6
DOIs
StatePublished - Apr 2025
Externally publishedYes

Keywords

  • Hierarchical reinforcement learning
  • Self-play training
  • Table tennis

Fingerprint

Dive into the research topics of 'Mastering table tennis with hierarchy: a reinforcement learning approach with progressive self-play training'. Together they form a unique fingerprint.

Cite this