TY - GEN
T1 - Learning from Hindsight Demonstrations
AU - Shao, Mengxuan
AU - Jiang, Feng
AU - Liu, Shaohui
AU - Han, Kun
AU - Zhao, Debin
N1 - Publisher Copyright:
© 2023, The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd.
PY - 2023
Y1 - 2023
N2 - Learning from demonstrations (LfD) is an important technique to help reinforcement learning (RL) boost the training process, especially in the case of sparse rewards. But a major obstacle is the acquisition of expert demonstrations, which is difficult or expensive to obtain in many cases. In this paper, we propose a unique method called Learning from Hindsight Demonstrations (LfHD) to automatically produce hindsight demonstrations, on which LfD can be performed and the cost of acquiring expert demonstrations is avoided. The produced demonstrations are comparable to those of experts at certain success rate. We also improve the LfD method to make better use of the produced demonstrations. Experiments show that our method can greatly improve the training efficiency compared to existing algorithms.
AB - Learning from demonstrations (LfD) is an important technique to help reinforcement learning (RL) boost the training process, especially in the case of sparse rewards. But a major obstacle is the acquisition of expert demonstrations, which is difficult or expensive to obtain in many cases. In this paper, we propose a unique method called Learning from Hindsight Demonstrations (LfHD) to automatically produce hindsight demonstrations, on which LfD can be performed and the cost of acquiring expert demonstrations is avoided. The produced demonstrations are comparable to those of experts at certain success rate. We also improve the LfD method to make better use of the produced demonstrations. Experiments show that our method can greatly improve the training efficiency compared to existing algorithms.
KW - hindsight experience replay
KW - learning from demonstrations
KW - reinforcement learning
UR - https://www.scopus.com/pages/publications/85161641195
U2 - 10.1007/978-981-99-1642-9_41
DO - 10.1007/978-981-99-1642-9_41
M3 - 会议稿件
AN - SCOPUS:85161641195
SN - 9789819916412
T3 - Communications in Computer and Information Science
SP - 480
EP - 491
BT - Neural Information Processing - 29th International Conference, ICONIP 2022, Proceedings
A2 - Tanveer, Mohammad
A2 - Agarwal, Sonali
A2 - Ozawa, Seiichi
A2 - Ekbal, Asif
A2 - Jatowt, Adam
PB - Springer Science and Business Media Deutschland GmbH
T2 - 29th International Conference on Neural Information Processing, ICONIP 2022
Y2 - 22 November 2022 through 26 November 2022
ER -