Skip to main navigation Skip to search Skip to main content

Reinforcement learning for control with probabilistic stability guarantee: A finite-sample approach

  • Minghao Han
  • , Lixian Zhang*
  • , Chenliang Liu
  • , Zhipeng Zhou
  • , Jun Wang
  • , Wei Pan
  • *Corresponding author for this work
  • Department of Control Science and Engineering
  • Harbin Institute of Technology
  • Schools of Automation
  • Central South University
  • Delft University of Technology
  • University College London
  • Department of Computer Science
  • University of Manchester

Research output: Contribution to journalArticlepeer-review

Abstract

This paper presents a novel approach to reinforcement learning (RL) for control systems that provides probabilistic stability guarantees using finite data. Leveraging Lyapunov's method, we propose a probabilistic stability theorem that ensures mean square stability using only a finite number of sampled trajectories. The probability of stability increases with the number and length of trajectories, converging to certainty as data size grows. Additionally, we derive a policy gradient theorem for stabilizing policy learning and develop an RL algorithm, L-REINFORCE, that extends the classical REINFORCE algorithm to stabilization problems. The effectiveness of L-REINFORCE is demonstrated through simulations on a Cartpole task, where it outperforms the baseline in ensuring stability. This work bridges a critical gap between RL and control theory, enabling stability analysis and controller design in a model-free framework with finite data.

Original languageEnglish
Article number112964
JournalAutomatica
Volume188
DOIs
StatePublished - Jun 2026

Keywords

  • Finite sample
  • Lyapunov's method
  • Nonlinear control
  • Probabilistic bound
  • Reinforcement learning

Fingerprint

Dive into the research topics of 'Reinforcement learning for control with probabilistic stability guarantee: A finite-sample approach'. Together they form a unique fingerprint.

Cite this