Skip to main navigation Skip to search Skip to main content

Robustness-Guaranteed Reinforcement Learning under Uncertainties in Dynamics Modeling and State Estimates

  • Duofeng Pan
  • , Yongquan Huang
  • , Rong Chen
  • , Wenjie Lu*
  • , Manman Hu
  • *Corresponding author for this work
  • School of Robotics and Advanced Manufacture, Harbin Institute of Technology Shenzhen
  • The University of Hong Kong

Research output: Contribution to journalArticlepeer-review

Abstract

Robust reinforcement learning enhances system robustness against uncertainties through bilevel optimization. However, the controller learned by reinforcement learning struggles to offer strict robustness guarantee, limiting its reliability in realworld applications. This paper presents robustness-guaranteed reinforcement learning framework under bounded uncertainties in dynamics modeling and state estimates. By describing the uncertainties and system dynamics models via ReLU neural networks, the states that most violate the proposed robustness conditions can be exactly found, and their robustness is prioritized during training. The solvability of the reinforcement learning problem is particularly studied, yielding a necessary condition that captures the adversarial relationship between control saturation and uncertainties. Subsequently, the attractive region can be pre-validated and the maximum robustness of the controller can be achieved by amplifying the uncertainties to a derived maximum extent in training. The proposed method has been fully evaluated in simulation and its advantages have been proven, including 3D quadrotor reference trajectory tracking tasks.

Original languageEnglish
JournalIEEE Transactions on Artificial Intelligence
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Guaranteed Robustness
  • Piecewise Linear
  • Robust Reinforcement Learning
  • Uncertain System

Fingerprint

Dive into the research topics of 'Robustness-Guaranteed Reinforcement Learning under Uncertainties in Dynamics Modeling and State Estimates'. Together they form a unique fingerprint.

Cite this