Skip to main navigation Skip to search Skip to main content

Capturing Cross-Modal Semantics by Generating Comments for Image-Text Contents

  • Faculty of Computing, Harbin Institute of Technology
  • Tencent

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent studies on multi-modal learning have benefited much from the development of Vision-Language Pre-training (VLP) approaches, which are believed to be able to bridge the semantic gap between the visual and linguistic modalities. Despite the notable successes, the current VLP methodologies mainly focus on the semantic overlap of the visual and textual inputs, and thus the actual cross-modal information complementarity tends to be ignored. Consequently, this paper defines a new VLP task of generating natural-language comments from the given Image-Text contents, named as Cross-Modal Information Complementarity based Comment Generation (CroMIC-CMT), so as to guide the models to capture the cross-modal semantics. A generic architecture is also presented for the proposed task, with formalized components for the bidirectional multi-modal encoding and autoregressive text generation. Furthermore, to validate the effectiveness of CroMIC-CMT as a pre-training method, it is evaluated and compared on the downstream tasks. The experimental results show that the proposed VLP task brings a new perspective on multi-modal semantic learning, and can be taken as a potential pre-training paradigm to address the downstream problems 1The source code of this work is uploaded (here).

Original languageEnglish
Title of host publicationPattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
EditorsJosef Kittler, Hongkai Xiong, Jian Yang, Xilin Chen, Jiwen Lu, Weiyao Lin, Jingyi Yu, Weishi Zheng
PublisherSpringer Science and Business Media Deutschland GmbH
Pages135-148
Number of pages14
ISBN (Print)9789819556786
DOIs
StatePublished - 2026
Externally publishedYes
Event8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025 - Shanghai, China
Duration: 15 Oct 202518 Oct 2025

Publication series

NameLecture Notes in Computer Science
Volume16277 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
Country/TerritoryChina
CityShanghai
Period15/10/2518/10/25

Keywords

  • Comment Generation
  • Information Complementarity
  • Vision-Language Pre-training

Fingerprint

Dive into the research topics of 'Capturing Cross-Modal Semantics by Generating Comments for Image-Text Contents'. Together they form a unique fingerprint.

Cite this