Skip to main navigation Skip to search Skip to main content

A Model-Based Exploration Policy in Deep Q-Network

  • Shuailong Li
  • , Wei Zhang*
  • , Yuquan Leng*
  • , Xin Zhang
  • *Corresponding author for this work
  • University of Chinese Academy of Sciences
  • Southern University of Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Reinforcement learning has successfully been used in many applications and achieved prodigious performance (such as video games), and DQN is a well-known algorithm in RL. However, there are some disadvantages in practical applications, and the exploration and exploitation dilemma is one of them. To solve this problem, common strategies about exploration like ϵ-greedy have risen. Unfortunately, there are sample inefficient and ineffective because of the uncertainty of later exploration. In this paper, we propose a model-based exploration method that learns the state transition model to explore. Using the training rules of machine learning, we can train the state transition model networks to improve exploration efficiency and sample efficiency. We compare our algorithm with ϵ-greedy on the Deep Q-Networks (DQN) algorithm and apply it to the Atari 2600 games. Our algorithm outperforms the decaying ϵ-greedy strategy when we evaluate our algorithm across 14 Atari games in the Arcade Learning Environment (ALE).

Original languageEnglish
Title of host publication2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages336-343
Number of pages8
ISBN (Electronic)9781665406307
DOIs
StatePublished - 2021
Externally publishedYes
Event2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021 - Virtual, Chengdu, China
Duration: 3 Dec 20214 Dec 2021

Publication series

Name2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021

Conference

Conference2021 International Conference on Digital Society and Intelligent Systems, DSInS 2021
Country/TerritoryChina
CityVirtual, Chengdu
Period3/12/214/12/21

Keywords

  • exploration and exploitation dilemma
  • model-based exploration method
  • reinforcement learning

Fingerprint

Dive into the research topics of 'A Model-Based Exploration Policy in Deep Q-Network'. Together they form a unique fingerprint.

Cite this