TY - GEN
T1 - Knowledge-Driven Actor-Critic VLMs for Autonomous Driving in Adverse Weather Condition
AU - Yu, Jingqi
AU - Gao, Yongji
AU - Zou, Qiming
AU - Mei, Yu
AU - Wang, Ling
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Vision Language Models (VLMs) significantly enhance scene understanding and planning capabilities for autonomous driving, yet they are prone to hallucinations. We have developed an actor-critic framework to mitigate these hallucinations. Specifically, the VLM serves as the actor, processing visual observations as input and generating a sequence of actions. A Large Language Model (LLM) acts as the critic, providing decision quality scores and corresponding textual explanations. These scores are calculated across various dimensions, including reasonableness, correctness, and conciseness. Based on the textual explanations, the LLM selects a prompt from an external knowl- edge base to guide the VLM in generating an action sequence that could improve the quality scores. The interaction between the actor and critic is expected to reduce hallucinations due to the utilization of the external knowledge base. To evaluate the performance of our framework under challenging driving conditions, we conducted experiments on a real-world dataset where images were collected in snowy scenarios. The experimental results indicate that, with increasing rounds of interaction between the actor and critic, the correctness and reasonableness scores for tasks increased by 45% compared to the standard VLM.
AB - Vision Language Models (VLMs) significantly enhance scene understanding and planning capabilities for autonomous driving, yet they are prone to hallucinations. We have developed an actor-critic framework to mitigate these hallucinations. Specifically, the VLM serves as the actor, processing visual observations as input and generating a sequence of actions. A Large Language Model (LLM) acts as the critic, providing decision quality scores and corresponding textual explanations. These scores are calculated across various dimensions, including reasonableness, correctness, and conciseness. Based on the textual explanations, the LLM selects a prompt from an external knowl- edge base to guide the VLM in generating an action sequence that could improve the quality scores. The interaction between the actor and critic is expected to reduce hallucinations due to the utilization of the external knowledge base. To evaluate the performance of our framework under challenging driving conditions, we conducted experiments on a real-world dataset where images were collected in snowy scenarios. The experimental results indicate that, with increasing rounds of interaction between the actor and critic, the correctness and reasonableness scores for tasks increased by 45% compared to the standard VLM.
KW - Actor-Critic
KW - Autonomous Driving
KW - RAG
KW - VLM
UR - https://www.scopus.com/pages/publications/105043739455
U2 - 10.1109/AIITA69518.2026.11567065
DO - 10.1109/AIITA69518.2026.11567065
M3 - 会议稿件
AN - SCOPUS:105043739455
T3 - 2026 6th International Conference on Artificial Intelligence and Industrial Technology Applications, AIITA 2026
SP - 598
EP - 602
BT - 2026 6th International Conference on Artificial Intelligence and Industrial Technology Applications, AIITA 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 6th International Conference on Artificial Intelligence and Industrial Technology Applications, AIITA 2026
Y2 - 10 April 2026 through 12 April 2026
ER -