Abstract
The multi-modal Language Models (LMs) perform very well on alignment-style tasks such as Image–Text Retrieval and Image Captioning, benefiting mainly from pre-training on numerous image–text pairs. However, our evaluations indicate that these models underperform on conversation-style multi-modal tasks, such as Image-Chat and Visual Dialog, which constitute a crucial segment of multi-modal applications. To bridge this gap, this paper proposes a novel pre-training task, named as MBCG, to stimulate the abilities of existing multi-modal LMs on conversation-style multi-modal tasks without hurting their intrinsic abilities. For this purpose, we collect two image–text-comments triplet multi-modal datasets in both English and Chinese to apply the new pre-training task to existing models. The experimental results reveal that the MBCG task can significantly boost the performance of these models on conversation-style tasks, without any noticeable performance decline on their original evaluation tasks.
| Original language | English |
|---|---|
| Article number | 103047 |
| Journal | Information Fusion |
| Volume | 120 |
| DOIs | |
| State | Published - Aug 2025 |
Keywords
- Conversation-style tasks
- Multi-modal LMs
- New dataset
Fingerprint
Dive into the research topics of 'Stimulating conversation-style emergencies of multi-modal LMs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver