TY - GEN
T1 - Average reward reinforcement learning for semi-markov decision processes
AU - Yang, Jiayuan
AU - Li, Yanjie
AU - Chen, Haoyao
AU - Li, Jiangang
N1 - Publisher Copyright:
© Springer International Publishing AG 2017.
PY - 2017
Y1 - 2017
N2 - In this paper, we study new reinforcement learning (RL) algorithms for Semi-Markov decision processes (SMDPs) with an average reward criterion. Based on the discrete-time type Bellman optimality equation, we use incremental value iteration (IVI), stochastic shortest path (SSP) value iteration and bisection algorithms to derive novel RL algorithms in a straightforward way. These algorithms use IVI, SSP and dichotomy to directly estimate the optimal average reward to solve the instability of average reward RL, respectively. Furthermore, a simulation experiment is used to compare the convergence among these algorithms.
AB - In this paper, we study new reinforcement learning (RL) algorithms for Semi-Markov decision processes (SMDPs) with an average reward criterion. Based on the discrete-time type Bellman optimality equation, we use incremental value iteration (IVI), stochastic shortest path (SSP) value iteration and bisection algorithms to derive novel RL algorithms in a straightforward way. These algorithms use IVI, SSP and dichotomy to directly estimate the optimal average reward to solve the instability of average reward RL, respectively. Furthermore, a simulation experiment is used to compare the convergence among these algorithms.
KW - Incremental value iteration
KW - SMDPs
KW - Stochastic shortest path
UR - https://www.scopus.com/pages/publications/85035143316
U2 - 10.1007/978-3-319-70087-8_79
DO - 10.1007/978-3-319-70087-8_79
M3 - 会议稿件
AN - SCOPUS:85035143316
SN - 9783319700861
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 768
EP - 777
BT - Neural Information Processing - 24th International Conference, ICONIP 2017, Proceedings
A2 - Li, Yuanqing
A2 - Liu, Derong
A2 - Xie, Shengli
A2 - El-Alfy, El-Sayed M.
A2 - Zhao, Dongbin
PB - Springer Verlag
T2 - 24th International Conference on Neural Information Processing, ICONIP 2017
Y2 - 14 November 2017 through 18 November 2017
ER -