TY - GEN
T1 - Enhancing Compositional Reasoning in Multimodal Large Language Models
AU - Qian, Shun
AU - Liu, Bingquan
AU - Sun, Chengjie
AU - Xu, Zhen
AU - Wang, Baoxun
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Recent studies have highlighted the shortage in compositional reasoning capabilities within Vision-Language Models (VLMs) such as CLIP and SigLIP. While these deficiencies have been documented in VLMs, their transmission to Multimodal Large Language Models (MLLMs)—which frequently rely on VLM-derived visual encoders—has not been thoroughly explored. This paper systematically investigates the compositional reasoning deficiencies in MLLMs and demonstrates that these models inherit the limitations of their underlying visual architectures. To address this challenge, we propose a vision-language compositionality enhancement strategy aimed at improving compositional reasoning while maintaining the MLLM’s overall multimodal understanding and reasoning capabilities. Our approach integrates two key innovations: (1) adopting multi-layer visual features as inputs of MLLMs to enhance fine-grained visual information, which is essential for effective compositional reasoning, and (2) a contrastive learning (CL) stage intended to improve the compositional reasoning abilities of MLLMs. Extensive experimental evaluations demonstrate that our auxiliary training strategy significantly enhances the performance of LLaVA-1.5-7B on compositional reasoning benchmarks, achieving performance parity with GPT-4V on specific tasks. Importantly, this strategy not only preserves but also improves performance across general multimodal benchmarks, highlighting its dual efficacy in both specialization and generalization.
AB - Recent studies have highlighted the shortage in compositional reasoning capabilities within Vision-Language Models (VLMs) such as CLIP and SigLIP. While these deficiencies have been documented in VLMs, their transmission to Multimodal Large Language Models (MLLMs)—which frequently rely on VLM-derived visual encoders—has not been thoroughly explored. This paper systematically investigates the compositional reasoning deficiencies in MLLMs and demonstrates that these models inherit the limitations of their underlying visual architectures. To address this challenge, we propose a vision-language compositionality enhancement strategy aimed at improving compositional reasoning while maintaining the MLLM’s overall multimodal understanding and reasoning capabilities. Our approach integrates two key innovations: (1) adopting multi-layer visual features as inputs of MLLMs to enhance fine-grained visual information, which is essential for effective compositional reasoning, and (2) a contrastive learning (CL) stage intended to improve the compositional reasoning abilities of MLLMs. Extensive experimental evaluations demonstrate that our auxiliary training strategy significantly enhances the performance of LLaVA-1.5-7B on compositional reasoning benchmarks, achieving performance parity with GPT-4V on specific tasks. Importantly, this strategy not only preserves but also improves performance across general multimodal benchmarks, highlighting its dual efficacy in both specialization and generalization.
KW - Compositional Reasoning
KW - Multimodal Large Language Models
KW - Vision-Language Models
UR - https://www.scopus.com/pages/publications/105028432106
U2 - 10.1007/978-981-95-5679-3_6
DO - 10.1007/978-981-95-5679-3_6
M3 - 会议稿件
AN - SCOPUS:105028432106
SN - 9789819556786
T3 - Lecture Notes in Computer Science
SP - 76
EP - 90
BT - Pattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
A2 - Kittler, Josef
A2 - Xiong, Hongkai
A2 - Yang, Jian
A2 - Chen, Xilin
A2 - Lu, Jiwen
A2 - Lin, Weiyao
A2 - Yu, Jingyi
A2 - Zheng, Weishi
PB - Springer Science and Business Media Deutschland GmbH
T2 - 8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
Y2 - 15 October 2025 through 18 October 2025
ER -