Skip to main navigation Skip to search Skip to main content

Safe Reinforcement Learning in Autonomous Driving With Epistemic Uncertainty Estimation

  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Safety is one of the critical challenges in the autonomous driving task. Recent works address the safety by implementing a safe reinforcement learning (safe RL) mechanism. However, most approaches make conservative decisions without knowing the confidence of the actions, which ultimately causes traffic congestion and low travel efficiency. This paper proposes an uncertainty-augmented Lagrangian safe reinforcement algorithm (Lag-U) to improve exploration and safety performance for autonomous driving. First, epistemic uncertainty is introduced into safe RL by using deep ensemble. We use the estimated epistemic uncertainty to encourage exploration and to learn a risk-sensitive policy by adaptively modifying safety constraints. Second, we facilitate an intervention assurance to choose safer actions based on the quantified epistemic uncertainty during deployment. Experimental results prove that the proposed method outperforms other safe RL baselines. The trained vehicle can make a decent trade-off between high efficiency and avoiding risks, thus preventing ultra-conservative policy.

Original languageEnglish
Pages (from-to)13653-13666
Number of pages14
JournalIEEE Transactions on Intelligent Transportation Systems
Volume25
Issue number10
DOIs
StatePublished - 2024
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure

Keywords

  • Safe reinforcement learning
  • exploration
  • risk aware policy
  • uncertainty quantification

Fingerprint

Dive into the research topics of 'Safe Reinforcement Learning in Autonomous Driving With Epistemic Uncertainty Estimation'. Together they form a unique fingerprint.

Cite this