Skip to main navigation Skip to search Skip to main content

Man-Machine Collaborative Task Planning Based on Visual Language Models

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Traditional robot planning systems frequently exhibit limited semantic adaptability in unstructured environments and dynamic human-robot collaboration scenarios. To overcome this limitation, this paper proposes a hierarchical collaborative task planning framework based on visual language models (VLMs). Specifically, we use the Open X-Embodiment dataset to perform parameter efficient fine-tuning (PEFT) on the Open Flamingo model to align the visual language representation with the robot operation space. To enhance robustness against environmental and interaction uncertainties inherent in collaborative scenarios, we introduce a progressive react prompting (PRP) mechanism. This mechanism implements a structured three-level correction strategy, including local, sub-task, and global adjustments, to monitor the task execution status in real-time and correct errors. Empirical evaluations on the CALVIN benchmark show that our proposed method significantly outperforms the baseline method in terms of long-horizon task success rate. Additionally, practical experiments conducted on the Franka Emika arm verified the feasibility of this framework in complex collaborative environments.

Original languageEnglish
Title of host publicationProceedings of the 5th Conference on Fully Actuated System Theory and Applications, FASTA 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1418-1422
Number of pages5
ISBN (Electronic)9798319547323
DOIs
StatePublished - 2026
Event5th Conference on Fully Actuated System Theory and Applications, FASTA 2026 - Qinhuangdao, China
Duration: 22 May 202624 May 2026

Publication series

NameProceedings of the 5th Conference on Fully Actuated System Theory and Applications, FASTA 2026

Conference

Conference5th Conference on Fully Actuated System Theory and Applications, FASTA 2026
Country/TerritoryChina
CityQinhuangdao
Period22/05/2624/05/26

Keywords

  • Embodied AI
  • Human-Machine Collaboration
  • Task Planning
  • Visual Language Models

Fingerprint

Dive into the research topics of 'Man-Machine Collaborative Task Planning Based on Visual Language Models'. Together they form a unique fingerprint.

Cite this