Skip to main navigation Skip to search Skip to main content

TAG: Teacher-Advice Mechanism with Gaussian Process for Reinforcement Learning

  • Ke Lin
  • , Duantengchuan Li
  • , Yanjie Li*
  • , Shiyu Chen
  • , Qi Liu
  • , Jianqi Gao
  • , Yanrui Jin
  • , Liang Gong
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Wuhan University
  • Shanghai Jiao Tong University

Research output: Contribution to journalArticlepeer-review

Abstract

Reinforcement learning (RL) still suffers from the problem of sample inefficiency and struggles with the exploration issue, particularly in situations with long-delayed rewards, sparse rewards, and deep local optimum. Recently, learning from demonstration (LfD) paradigm was proposed to tackle this problem. However, these methods usually require a large number of demonstrations. In this study, we present a sample efficient teacher-advice mechanism with Gaussian process (TAG) by leveraging a few expert demonstrations. In TAG, a teacher model is built to provide both an advice action and its associated confidence value. Then, a guided policy is formulated to guide the agent in the exploration phase via the defined criteria. Through the TAG mechanism, the agent is capable of exploring the environment more intentionally. Moreover, with the confidence value, the guided policy can guide the agent precisely. Also, due to the strong generalization ability of Gaussian process, the teacher model can utilize the demonstrations more effectively. Therefore, substantial improvement in performance and sample efficiency can be attained. Considerable experiments on sparse reward environments demonstrate that the TAG mechanism can help typical RL algorithms achieve significant performance gains. In addition, the TAG mechanism with soft actor-critic algorithm (TAG-SAC) attains the state-of-the-art performance over other LfD counterparts on several delayed reward and complicated continuous control environments.

Original languageEnglish
Pages (from-to)12419-12433
Number of pages15
JournalIEEE Transactions on Neural Networks and Learning Systems
Volume35
Issue number9
DOIs
StatePublished - 2024
Externally publishedYes

Keywords

  • Gaussian process
  • reinforcement learning (RL)
  • teacher-advice mechanism

Fingerprint

Dive into the research topics of 'TAG: Teacher-Advice Mechanism with Gaussian Process for Reinforcement Learning'. Together they form a unique fingerprint.

Cite this