Abstract
The vision-language projector is a critical component in Multimodal Large Language Models (MLLMs), but its output of long visual token sequences poses a significant computational challenge. While previous studies have focused on shortening these sequences for efficiency, they have largely overlooked the inherent spatial discrepancy between a serialized 2D visual sequence and the 1D nature of language tokens, thereby sacrificing performance for efficiency. To tackle these challenges, we introduce the Spatial-Aware Efficient Projector (SAEP). Our method first restores multi-layer visual features into their original 2D spatial layout and then employs a novel convolution module to perform compression. This design allows SAEP to compress visual tokens by 75% (e.g., from 576 to 144), while outperforming existing state-of-the-art efficient projectors across general multimodal benchmarks. Crucially, SAEP excels in tasks demanding detailed spatial reasoning, proving its ability to preserve both local visual details and global semantics within a highly compact visual token sequence. 11The source code is publicly released at https://github.com/AAbathur/SAEP-Projector.
| Original language | English |
|---|---|
| Article number | 132839 |
| Journal | Expert Systems with Applications |
| Volume | 328 |
| DOIs | |
| State | Published - 1 Oct 2026 |
Keywords
- Efficient projector
- Multi-layer feature aggregation
- Multi-modal large language models
- Spatial-aware
Fingerprint
Dive into the research topics of 'Spatial -aware efficient projector for MLLMs via multi-layer feature aggregation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver