Skip to main navigation Skip to search Skip to main content

Average reward reinforcement learning for semi-markov decision processes

  • Harbin Institute of Technology Shenzhen

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, we study new reinforcement learning (RL) algorithms for Semi-Markov decision processes (SMDPs) with an average reward criterion. Based on the discrete-time type Bellman optimality equation, we use incremental value iteration (IVI), stochastic shortest path (SSP) value iteration and bisection algorithms to derive novel RL algorithms in a straightforward way. These algorithms use IVI, SSP and dichotomy to directly estimate the optimal average reward to solve the instability of average reward RL, respectively. Furthermore, a simulation experiment is used to compare the convergence among these algorithms.

Original languageEnglish
Title of host publicationNeural Information Processing - 24th International Conference, ICONIP 2017, Proceedings
EditorsYuanqing Li, Derong Liu, Shengli Xie, El-Sayed M. El-Alfy, Dongbin Zhao
PublisherSpringer Verlag
Pages768-777
Number of pages10
ISBN (Print)9783319700861
DOIs
StatePublished - 2017
Externally publishedYes
Event24th International Conference on Neural Information Processing, ICONIP 2017 - Guangzhou, China
Duration: 14 Nov 201718 Nov 2017

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume10634 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference24th International Conference on Neural Information Processing, ICONIP 2017
Country/TerritoryChina
CityGuangzhou
Period14/11/1718/11/17

Keywords

  • Incremental value iteration
  • SMDPs
  • Stochastic shortest path

Fingerprint

Dive into the research topics of 'Average reward reinforcement learning for semi-markov decision processes'. Together they form a unique fingerprint.

Cite this