Skip to main navigation Skip to search Skip to main content

Stimulating conversation-style emergencies of multi-modal LMs

  • Harbin Institute of Technology
  • Tencent

Research output: Contribution to journalArticlepeer-review

Abstract

The multi-modal Language Models (LMs) perform very well on alignment-style tasks such as Image–Text Retrieval and Image Captioning, benefiting mainly from pre-training on numerous image–text pairs. However, our evaluations indicate that these models underperform on conversation-style multi-modal tasks, such as Image-Chat and Visual Dialog, which constitute a crucial segment of multi-modal applications. To bridge this gap, this paper proposes a novel pre-training task, named as MBCG, to stimulate the abilities of existing multi-modal LMs on conversation-style multi-modal tasks without hurting their intrinsic abilities. For this purpose, we collect two image–text-comments triplet multi-modal datasets in both English and Chinese to apply the new pre-training task to existing models. The experimental results reveal that the MBCG task can significantly boost the performance of these models on conversation-style tasks, without any noticeable performance decline on their original evaluation tasks.

Original languageEnglish
Article number103047
JournalInformation Fusion
Volume120
DOIs
StatePublished - Aug 2025

Keywords

  • Conversation-style tasks
  • Multi-modal LMs
  • New dataset

Fingerprint

Dive into the research topics of 'Stimulating conversation-style emergencies of multi-modal LMs'. Together they form a unique fingerprint.

Cite this