TY - GEN
T1 - Deep Reinforcement Learning Approaches for Motion Planning of Autonomous Buses in Uncertain Environments
AU - Zhao, Hantao
AU - Hu, Tai
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Advancements in sensor and communication technologies have accelerated the development of autonomous driving, expanding its application possibilities. The primary challenge is the uncertainty in dynamic urban environments, which affects motion planning algorithms. This paper develops a partially observable markov decision framework tailored for autonomous buses, integrating behavioral decision-making with motion planning. First, we propose an environmental modeling method sensitive to uncertain factors and a specialized hardware framework for autonomous buses. Second, we derived a mathematical model for optimal bus centering on curved roads, informing the reward structure for reinforcement learning. Third, utilizing historical data, we implemented a time-dependent deep reinforcement learning algorithm, recursive deterministic policy gradient (RDPG), to enhance observation accuracy and determine the optimal driving strategy. Our simulations confirm that this algorithm surpasses existing technologies in performance.
AB - Advancements in sensor and communication technologies have accelerated the development of autonomous driving, expanding its application possibilities. The primary challenge is the uncertainty in dynamic urban environments, which affects motion planning algorithms. This paper develops a partially observable markov decision framework tailored for autonomous buses, integrating behavioral decision-making with motion planning. First, we propose an environmental modeling method sensitive to uncertain factors and a specialized hardware framework for autonomous buses. Second, we derived a mathematical model for optimal bus centering on curved roads, informing the reward structure for reinforcement learning. Third, utilizing historical data, we implemented a time-dependent deep reinforcement learning algorithm, recursive deterministic policy gradient (RDPG), to enhance observation accuracy and determine the optimal driving strategy. Our simulations confirm that this algorithm surpasses existing technologies in performance.
KW - autonomous buses
KW - motion planning
KW - partial observed Markov decision process (POMDP)
KW - recursive deterministic policy gradient (RDPG)
UR - https://www.scopus.com/pages/publications/105043181204
U2 - 10.1007/978-981-95-8620-2_10
DO - 10.1007/978-981-95-8620-2_10
M3 - 会议稿件
AN - SCOPUS:105043181204
SN - 9789819586196
T3 - Lecture Notes in Electrical Engineering
SP - 142
EP - 154
BT - Resilience Transportation and Mobility Safety
A2 - Wang, Wuhong
A2 - Ci, Yusheng
A2 - Hu, Xiaowei
A2 - Tan, Haiqiu
A2 - Li, Min
PB - Springer Science and Business Media Deutschland GmbH
T2 - 16th International Conference on Green Intelligent Transportation System and Safety, GITSS 2025
Y2 - 9 May 2025 through 11 May 2025
ER -