Skip to main navigation Skip to search Skip to main content

Spatial -aware efficient projector for MLLMs via multi-layer feature aggregation

  • Harbin Institute of Technology
  • Tencent PCG Application Research Center

Research output: Contribution to journalArticlepeer-review

Abstract

The vision-language projector is a critical component in Multimodal Large Language Models (MLLMs), but its output of long visual token sequences poses a significant computational challenge. While previous studies have focused on shortening these sequences for efficiency, they have largely overlooked the inherent spatial discrepancy between a serialized 2D visual sequence and the 1D nature of language tokens, thereby sacrificing performance for efficiency. To tackle these challenges, we introduce the Spatial-Aware Efficient Projector (SAEP). Our method first restores multi-layer visual features into their original 2D spatial layout and then employs a novel convolution module to perform compression. This design allows SAEP to compress visual tokens by 75% (e.g., from 576 to 144), while outperforming existing state-of-the-art efficient projectors across general multimodal benchmarks. Crucially, SAEP excels in tasks demanding detailed spatial reasoning, proving its ability to preserve both local visual details and global semantics within a highly compact visual token sequence. 11The source code is publicly released at https://github.com/AAbathur/SAEP-Projector.

Original languageEnglish
Article number132839
JournalExpert Systems with Applications
Volume328
DOIs
StatePublished - 1 Oct 2026

Keywords

  • Efficient projector
  • Multi-layer feature aggregation
  • Multi-modal large language models
  • Spatial-aware

Fingerprint

Dive into the research topics of 'Spatial -aware efficient projector for MLLMs via multi-layer feature aggregation'. Together they form a unique fingerprint.

Cite this