TY - GEN
T1 - Man-Machine Collaborative Task Planning Based on Visual Language Models
AU - Liu, Xian
AU - Sun, Weichao
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Traditional robot planning systems frequently exhibit limited semantic adaptability in unstructured environments and dynamic human-robot collaboration scenarios. To overcome this limitation, this paper proposes a hierarchical collaborative task planning framework based on visual language models (VLMs). Specifically, we use the Open X-Embodiment dataset to perform parameter efficient fine-tuning (PEFT) on the Open Flamingo model to align the visual language representation with the robot operation space. To enhance robustness against environmental and interaction uncertainties inherent in collaborative scenarios, we introduce a progressive react prompting (PRP) mechanism. This mechanism implements a structured three-level correction strategy, including local, sub-task, and global adjustments, to monitor the task execution status in real-time and correct errors. Empirical evaluations on the CALVIN benchmark show that our proposed method significantly outperforms the baseline method in terms of long-horizon task success rate. Additionally, practical experiments conducted on the Franka Emika arm verified the feasibility of this framework in complex collaborative environments.
AB - Traditional robot planning systems frequently exhibit limited semantic adaptability in unstructured environments and dynamic human-robot collaboration scenarios. To overcome this limitation, this paper proposes a hierarchical collaborative task planning framework based on visual language models (VLMs). Specifically, we use the Open X-Embodiment dataset to perform parameter efficient fine-tuning (PEFT) on the Open Flamingo model to align the visual language representation with the robot operation space. To enhance robustness against environmental and interaction uncertainties inherent in collaborative scenarios, we introduce a progressive react prompting (PRP) mechanism. This mechanism implements a structured three-level correction strategy, including local, sub-task, and global adjustments, to monitor the task execution status in real-time and correct errors. Empirical evaluations on the CALVIN benchmark show that our proposed method significantly outperforms the baseline method in terms of long-horizon task success rate. Additionally, practical experiments conducted on the Franka Emika arm verified the feasibility of this framework in complex collaborative environments.
KW - Embodied AI
KW - Human-Machine Collaboration
KW - Task Planning
KW - Visual Language Models
UR - https://www.scopus.com/pages/publications/105043532580
U2 - 10.1109/FASTA70174.2026.11549542
DO - 10.1109/FASTA70174.2026.11549542
M3 - 会议稿件
AN - SCOPUS:105043532580
T3 - Proceedings of the 5th Conference on Fully Actuated System Theory and Applications, FASTA 2026
SP - 1418
EP - 1422
BT - Proceedings of the 5th Conference on Fully Actuated System Theory and Applications, FASTA 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 5th Conference on Fully Actuated System Theory and Applications, FASTA 2026
Y2 - 22 May 2026 through 24 May 2026
ER -