Skip to main navigation Skip to search Skip to main content

The policy gradient estimation of continuous-time Hidden Markov decision processes

  • Yanjie Li*
  • , Baoqun Yin
  • , Hongsheng Xi
  • *Corresponding author for this work
  • University of Science and Technology of China

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recently, gradient based methods have received much attention to optimize some dynamic systems with hidden information, such as routing problems of robotic systems. In this paper, we presented a process-Continuous Time Hidden Markov Decision Process (CTHMDP), which can be used to model the robotic systems. For this process, the problem of policy gradient estimation is studied. Firstly, an approximation formula to the gradient is presented, then by using the uniformization method, we introduce an algorithm, which can be considered as an extension of Gradient of Partially Observable Markov Decision Process (GPOMDP) algorithm to the continue time model. Finally, the convergence and error bound of the algorithm are considered.

Original languageEnglish
Title of host publicationICIA 2005 - Proceedings of 2005 International Conference on Information Acquisition
Pages300-303
Number of pages4
StatePublished - 2005
Externally publishedYes
EventICIA 2005: 2005 International Conference on Information Acquisition - Hong Kong, China
Duration: 27 Jun 20053 Jul 2005

Publication series

NameICIA 2005 - Proceedings of 2005 International Conference on Information Acquisition
Volume2005

Conference

ConferenceICIA 2005: 2005 International Conference on Information Acquisition
Country/TerritoryChina
CityHong Kong
Period27/06/053/07/05

Fingerprint

Dive into the research topics of 'The policy gradient estimation of continuous-time Hidden Markov decision processes'. Together they form a unique fingerprint.

Cite this