TY - GEN
T1 - Language-Conditioned Open-Vocabulary Mobile Manipulation with Pretrained Models
AU - Tan, Shen
AU - Zhou, Dong
AU - Shao, Xiangyu
AU - Wang, Junqiao
AU - Sun, Guanghui
N1 - Publisher Copyright:
© 2025 International Joint Conferences on Artificial Intelligence. All rights reserved.
PY - 2025
Y1 - 2025
N2 - Open-vocabulary mobile manipulation (OVMM) that involves the handling of novel and unseen objects across different workspaces remains a significant challenge for real-world robotic applications. In this paper, we propose a novel Language-conditioned Open-Vocabulary Mobile Manipulation framework, named LOVMM, incorporating the large language model (LLM) and vision-language model (VLM) to tackle various mobile manipulation tasks in household environments. Our approach is capable of solving various OVMM tasks with free-form natural language instructions (e.g. “toss the food boxes on the office room desk to the trash bin in the corner”, and “pack the bottles from the bed to the box in the guestroom”). Extensive experiments simulated in complex household environments show strong zero-shot generalization and multi-task learning abilities of LOVMM. Moreover, our approach can also generalize to multiple tabletop manipulation tasks and achieve better success rates compared to other state-of-the-art methods.
AB - Open-vocabulary mobile manipulation (OVMM) that involves the handling of novel and unseen objects across different workspaces remains a significant challenge for real-world robotic applications. In this paper, we propose a novel Language-conditioned Open-Vocabulary Mobile Manipulation framework, named LOVMM, incorporating the large language model (LLM) and vision-language model (VLM) to tackle various mobile manipulation tasks in household environments. Our approach is capable of solving various OVMM tasks with free-form natural language instructions (e.g. “toss the food boxes on the office room desk to the trash bin in the corner”, and “pack the bottles from the bed to the box in the guestroom”). Extensive experiments simulated in complex household environments show strong zero-shot generalization and multi-task learning abilities of LOVMM. Moreover, our approach can also generalize to multiple tabletop manipulation tasks and achieve better success rates compared to other state-of-the-art methods.
UR - https://www.scopus.com/pages/publications/105021808319
U2 - 10.24963/ijcai.2025/976
DO - 10.24963/ijcai.2025/976
M3 - 会议稿件
AN - SCOPUS:105021808319
T3 - IJCAI International Joint Conference on Artificial Intelligence
SP - 8778
EP - 8786
BT - Proceedings of the 34th International Joint Conference on Artificial Intelligence, IJCAI 2025
A2 - Kwok, James
PB - International Joint Conferences on Artificial Intelligence
T2 - 34th Internationa Joint Conference on Artificial Intelligence, IJCAI 2025
Y2 - 16 August 2025 through 22 August 2025
ER -