Skip to main navigation Skip to search Skip to main content

Enhancing Compositional Reasoning in Multimodal Large Language Models

  • Faculty of Computing, Harbin Institute of Technology
  • Tencent

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent studies have highlighted the shortage in compositional reasoning capabilities within Vision-Language Models (VLMs) such as CLIP and SigLIP. While these deficiencies have been documented in VLMs, their transmission to Multimodal Large Language Models (MLLMs)—which frequently rely on VLM-derived visual encoders—has not been thoroughly explored. This paper systematically investigates the compositional reasoning deficiencies in MLLMs and demonstrates that these models inherit the limitations of their underlying visual architectures. To address this challenge, we propose a vision-language compositionality enhancement strategy aimed at improving compositional reasoning while maintaining the MLLM’s overall multimodal understanding and reasoning capabilities. Our approach integrates two key innovations: (1) adopting multi-layer visual features as inputs of MLLMs to enhance fine-grained visual information, which is essential for effective compositional reasoning, and (2) a contrastive learning (CL) stage intended to improve the compositional reasoning abilities of MLLMs. Extensive experimental evaluations demonstrate that our auxiliary training strategy significantly enhances the performance of LLaVA-1.5-7B on compositional reasoning benchmarks, achieving performance parity with GPT-4V on specific tasks. Importantly, this strategy not only preserves but also improves performance across general multimodal benchmarks, highlighting its dual efficacy in both specialization and generalization.

Original languageEnglish
Title of host publicationPattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
EditorsJosef Kittler, Hongkai Xiong, Jian Yang, Xilin Chen, Jiwen Lu, Weiyao Lin, Jingyi Yu, Weishi Zheng
PublisherSpringer Science and Business Media Deutschland GmbH
Pages76-90
Number of pages15
ISBN (Print)9789819556786
DOIs
StatePublished - 2026
Event8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025 - Shanghai, China
Duration: 15 Oct 202518 Oct 2025

Publication series

NameLecture Notes in Computer Science
Volume16277 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
Country/TerritoryChina
CityShanghai
Period15/10/2518/10/25

Keywords

  • Compositional Reasoning
  • Multimodal Large Language Models
  • Vision-Language Models

Fingerprint

Dive into the research topics of 'Enhancing Compositional Reasoning in Multimodal Large Language Models'. Together they form a unique fingerprint.

Cite this