Skip to main navigation Skip to search Skip to main content

Generalized Critic Policy Optimization: A Model for Combining Advantage Estimates in Actor Critic Methods

  • School of Computer Science and Technology, Harbin Institute of Technology
  • Frères Mentouri Constantine 1 University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We present a general model for actor critic methods that represent the possibility of combining value function estimations as a means to further reduce the policy gradient's variance and improve the learning result. We show the potential of this architecture by implementing an example case to learn some of the Pybullet continous control robotic tasks with OpenAI Gym. We show by experimenting with a special case the effect of the external parameters on the overall performance of the policy optimization algorithm.

Original languageEnglish
Title of host publication2020 IEEE International Conference on Image Processing, ICIP 2020 - Proceedings
PublisherIEEE Computer Society
Pages3184-3188
Number of pages5
ISBN (Electronic)9781728163956
DOIs
StatePublished - Oct 2020
Externally publishedYes
Event2020 IEEE International Conference on Image Processing, ICIP 2020 - Virtual, Abu Dhabi, United Arab Emirates
Duration: 25 Sep 202028 Sep 2020

Publication series

NameProceedings - International Conference on Image Processing, ICIP
Volume2020-October
ISSN (Print)1522-4880

Conference

Conference2020 IEEE International Conference on Image Processing, ICIP 2020
Country/TerritoryUnited Arab Emirates
CityVirtual, Abu Dhabi
Period25/09/2028/09/20

Keywords

  • actor critic
  • advantage estimation
  • deep reinforcement learning
  • policy gradient

Fingerprint

Dive into the research topics of 'Generalized Critic Policy Optimization: A Model for Combining Advantage Estimates in Actor Critic Methods'. Together they form a unique fingerprint.

Cite this