TY - GEN
T1 - Capturing Cross-Modal Semantics by Generating Comments for Image-Text Contents
AU - Qian, Shun
AU - Liu, Bingquan
AU - Sun, Chengjie
AU - Xu, Zhen
AU - Wang, Baoxun
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Recent studies on multi-modal learning have benefited much from the development of Vision-Language Pre-training (VLP) approaches, which are believed to be able to bridge the semantic gap between the visual and linguistic modalities. Despite the notable successes, the current VLP methodologies mainly focus on the semantic overlap of the visual and textual inputs, and thus the actual cross-modal information complementarity tends to be ignored. Consequently, this paper defines a new VLP task of generating natural-language comments from the given Image-Text contents, named as Cross-Modal Information Complementarity based Comment Generation (CroMIC-CMT), so as to guide the models to capture the cross-modal semantics. A generic architecture is also presented for the proposed task, with formalized components for the bidirectional multi-modal encoding and autoregressive text generation. Furthermore, to validate the effectiveness of CroMIC-CMT as a pre-training method, it is evaluated and compared on the downstream tasks. The experimental results show that the proposed VLP task brings a new perspective on multi-modal semantic learning, and can be taken as a potential pre-training paradigm to address the downstream problems 1The source code of this work is uploaded (here).
AB - Recent studies on multi-modal learning have benefited much from the development of Vision-Language Pre-training (VLP) approaches, which are believed to be able to bridge the semantic gap between the visual and linguistic modalities. Despite the notable successes, the current VLP methodologies mainly focus on the semantic overlap of the visual and textual inputs, and thus the actual cross-modal information complementarity tends to be ignored. Consequently, this paper defines a new VLP task of generating natural-language comments from the given Image-Text contents, named as Cross-Modal Information Complementarity based Comment Generation (CroMIC-CMT), so as to guide the models to capture the cross-modal semantics. A generic architecture is also presented for the proposed task, with formalized components for the bidirectional multi-modal encoding and autoregressive text generation. Furthermore, to validate the effectiveness of CroMIC-CMT as a pre-training method, it is evaluated and compared on the downstream tasks. The experimental results show that the proposed VLP task brings a new perspective on multi-modal semantic learning, and can be taken as a potential pre-training paradigm to address the downstream problems 1The source code of this work is uploaded (here).
KW - Comment Generation
KW - Information Complementarity
KW - Vision-Language Pre-training
UR - https://www.scopus.com/pages/publications/105028452199
U2 - 10.1007/978-981-95-5679-3_10
DO - 10.1007/978-981-95-5679-3_10
M3 - 会议稿件
AN - SCOPUS:105028452199
SN - 9789819556786
T3 - Lecture Notes in Computer Science
SP - 135
EP - 148
BT - Pattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
A2 - Kittler, Josef
A2 - Xiong, Hongkai
A2 - Yang, Jian
A2 - Chen, Xilin
A2 - Lu, Jiwen
A2 - Lin, Weiyao
A2 - Yu, Jingyi
A2 - Zheng, Weishi
PB - Springer Science and Business Media Deutschland GmbH
T2 - 8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
Y2 - 15 October 2025 through 18 October 2025
ER -