TY - GEN
T1 - LLM-Guided Exploration for Sample-Efficient UAV Navigation
AU - Xie, Xianan
AU - Li, Junbao
AU - Sheng, Yuanyuan
AU - Liu, Huanyu
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2027.
PY - 2027
Y1 - 2027
N2 - Autonomous navigation of unmanned aerial vehicles in unknown and cluttered environments remains challenging due to inefficient exploration and high sample complexity. While Large Language Models (LLMs) offer strong commonsense and spatial reasoning, they are not suited for real-time continuous control because of limited action precision and inference latency. To bridge this gap, we propose LLM-Guided Exploration, a hybrid training framework that leverages LLM reasoning to bootstrap the training of off-policy RL agents. Our approach utilizes the LLM as an initial supervisor, providing high-quality demonstrations to guide the agent through scenarios before gradually handing over control to the RL policy for fine-tuning. By decoupling high-level exploration reasoning from low-level motion execution, our framework enables structured and adaptive exploration without sacrificing control stability. Extensive experiments in diverse environments demonstrate that the proposed method significantly improves convergence speed, navigation success rate, and overall sample efficiency compared to existing exploration and navigation baselines.
AB - Autonomous navigation of unmanned aerial vehicles in unknown and cluttered environments remains challenging due to inefficient exploration and high sample complexity. While Large Language Models (LLMs) offer strong commonsense and spatial reasoning, they are not suited for real-time continuous control because of limited action precision and inference latency. To bridge this gap, we propose LLM-Guided Exploration, a hybrid training framework that leverages LLM reasoning to bootstrap the training of off-policy RL agents. Our approach utilizes the LLM as an initial supervisor, providing high-quality demonstrations to guide the agent through scenarios before gradually handing over control to the RL policy for fine-tuning. By decoupling high-level exploration reasoning from low-level motion execution, our framework enables structured and adaptive exploration without sacrificing control stability. Extensive experiments in diverse environments demonstrate that the proposed method significantly improves convergence speed, navigation success rate, and overall sample efficiency compared to existing exploration and navigation baselines.
KW - Deep Reinforcement Learning
KW - Large Language Models
KW - Unmanned Aerial Vehicles Navigation
UR - https://www.scopus.com/pages/publications/105047149788
U2 - 10.1007/978-3-032-31666-0_28
DO - 10.1007/978-3-032-31666-0_28
M3 - 会议稿件
AN - SCOPUS:105047149788
SN - 9783032316653
T3 - Lecture Notes in Computer Science
SP - 422
EP - 436
BT - Pattern Recognition - 28th International Conference, ICPR 2026, Proceedings
A2 - De Marsico, Maria
A2 - Ho, Tin Kam
A2 - Jurie, Frederic
A2 - Liu, Cheng-Lin
A2 - Lopresti, Daniel
A2 - Nyström, Ingela
A2 - Ogier, Jean-Marc
A2 - Ross, Arun
A2 - Wang, Liang
PB - Springer Science and Business Media Deutschland GmbH
T2 - 28th International Conference on Pattern Recognition, ICPR 2026
Y2 - 17 August 2026 through 22 August 2026
ER -