TY - GEN
T1 - A Model-Based Exploration Policy in Deep Q-Network
AU - Li, Shuailong
AU - Zhang, Wei
AU - Leng, Yuquan
AU - Zhang, Xin
N1 - Publisher Copyright:
© 2021 IEEE.
PY - 2021
Y1 - 2021
N2 - Reinforcement learning has successfully been used in many applications and achieved prodigious performance (such as video games), and DQN is a well-known algorithm in RL. However, there are some disadvantages in practical applications, and the exploration and exploitation dilemma is one of them. To solve this problem, common strategies about exploration like ϵ-greedy have risen. Unfortunately, there are sample inefficient and ineffective because of the uncertainty of later exploration. In this paper, we propose a model-based exploration method that learns the state transition model to explore. Using the training rules of machine learning, we can train the state transition model networks to improve exploration efficiency and sample efficiency. We compare our algorithm with ϵ-greedy on the Deep Q-Networks (DQN) algorithm and apply it to the Atari 2600 games. Our algorithm outperforms the decaying ϵ-greedy strategy when we evaluate our algorithm across 14 Atari games in the Arcade Learning Environment (ALE).
AB - Reinforcement learning has successfully been used in many applications and achieved prodigious performance (such as video games), and DQN is a well-known algorithm in RL. However, there are some disadvantages in practical applications, and the exploration and exploitation dilemma is one of them. To solve this problem, common strategies about exploration like ϵ-greedy have risen. Unfortunately, there are sample inefficient and ineffective because of the uncertainty of later exploration. In this paper, we propose a model-based exploration method that learns the state transition model to explore. Using the training rules of machine learning, we can train the state transition model networks to improve exploration efficiency and sample efficiency. We compare our algorithm with ϵ-greedy on the Deep Q-Networks (DQN) algorithm and apply it to the Atari 2600 games. Our algorithm outperforms the decaying ϵ-greedy strategy when we evaluate our algorithm across 14 Atari games in the Arcade Learning Environment (ALE).
KW - exploration and exploitation dilemma
KW - model-based exploration method
KW - reinforcement learning
UR - https://www.scopus.com/pages/publications/85125104191
U2 - 10.1109/DSInS54396.2021.9670573
DO - 10.1109/DSInS54396.2021.9670573
M3 - 会议稿件
AN - SCOPUS:85125104191
T3 - 2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021
SP - 336
EP - 343
BT - 2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021
Y2 - 3 December 2021 through 4 December 2021
ER -