Skip to main navigation Skip to search Skip to main content

Neural-Network-Driven Reward Prediction as a Heuristic: Advancing Q-Learning for Mobile Robot Path Planning

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

While Q-learning is a staple in path planning, it frequently suffers from slow convergence in practical scenarios. To overcome this, we propose NDR-QL (Neural-Network-Driven Reward Prediction based Q-learning), a method that leverages neural network outputs as heuristic information to accelerate the learning process. Specifically, we optimized a dual-output neural network by introducing a start-end channel separation mechanism and enhancing feature fusion. The resulting model generates two outputs: a narrowly focused "guideline"distribution and a broader "region"distribution. We utilize the guideline to calculate a continuous reward function and the region to initialize the Q-table with a strategic bias. Experiments on public datasets demonstrate that our model improves prediction accuracy by 5% over previous methods. Furthermore, NDR-QL accelerates convergence by 90% and produces superior path quality compared to existing improved Q-learning baselines.

Original languageEnglish
Title of host publication2026 12th International Conference on Control, Automation and Robotics, ICCAR 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages151-156
Number of pages6
Edition2026
ISBN (Electronic)9798319529350
DOIs
StatePublished - 2026
Event2026 12th International Conference on Control, Automation and Robotics, ICCAR 2026 - Nagoya, Japan
Duration: 8 Apr 202610 Apr 2026

Conference

Conference2026 12th International Conference on Control, Automation and Robotics, ICCAR 2026
Country/TerritoryJapan
CityNagoya
Period8/04/2610/04/26

Keywords

  • Neural network
  • Path planning
  • Q-learning algorithm
  • Reinforcement learning

Fingerprint

Dive into the research topics of 'Neural-Network-Driven Reward Prediction as a Heuristic: Advancing Q-Learning for Mobile Robot Path Planning'. Together they form a unique fingerprint.

Cite this