Skip to main navigation Skip to search Skip to main content

Adaptive dynamic programming for unknown continuous-time nonlinear systems via bias-policy iteration

  • Harbin Institute of Technology
  • National Key Laboratory of Complex System Control and Intelligent Agent Cooperation

Research output: Contribution to journalArticlepeer-review

Abstract

In this paper, a bias-policy iteration (Bias-PI) method is proposed to relax the requirement of the policy iteration method on the initial admissible control and achieve optimal control for unknown continuous-time nonlinear systems. First, a model-based Bias-PI method is introduced that uses a bias value function to ease the constraints of the initial admissible control. The boundedness of the bias value function and the convergence of the algorithm are demonstrated through rigorous mathematical proofs. Further, the data-driven implementation of the Bias-PI method is detailed, highlighting its ability to learn an optimal controller without prior system information, and simultaneously retaining the fast convergence properties of the traditional policy iteration algorithm. The effectiveness of the data-driven Bias-PI method is illustrated through two simulation examples.

Original languageEnglish
Article number112821
JournalAutomatica
Volume185
DOIs
StatePublished - Mar 2026

Keywords

  • Continuous-time nonlinear systems
  • Initial admissible control
  • Policy iteration
  • Reinforcement learning
  • Unknown dynamics

Fingerprint

Dive into the research topics of 'Adaptive dynamic programming for unknown continuous-time nonlinear systems via bias-policy iteration'. Together they form a unique fingerprint.

Cite this