Abstract
Vision-language models have advanced rapidly, yet they often rely on coarse, global image features and struggle with fine-grained cross-modal reasoning. The Visual Text Question Answering (VTQA) benchmark exemplifies this challenge: each instance pairs a natural image with a detailed text passage, requiring the alignment of specific entities across modalities for multi-hop reasoning. Existing approaches often employ frozen vision backbones that capture only global semantics, missing critical local details. To address this, we propose Multi-Scale Progressive Attention (MSPA). MSPA encodes images at multiple resolutions to obtain multi-scale features and applies a hierarchical top-k selection mechanism: at each scale, the model conditions on the textual query to identify the most relevant visual tokens, which are then projected to the next finer scale for refinement. This coarse-to-fine strategy concentrates computation on key regions, enabling detailed visual understanding without processing all high-resolution tiles. On the VTQA benchmark, MSPA establishes a new state-of-the-art, achieving 71.2% Exact Match (EM) on the test set (surpassing the previous best of 66.4%). These results demonstrate that integrating multi-scale visual features with progressive attention substantially enhances fine-grained visual-text reasoning, suggesting that coarse-to-fine processing can benefit other vision-language tasks demanding precise grounding.
| Original language | English |
|---|---|
| Article number | 134131 |
| Journal | Neurocomputing |
| Volume | 697 |
| DOIs | |
| State | Published - 7 Oct 2026 |
| Externally published | Yes |
Keywords
- Multi-scale attention
- Vision-language models
- Visual question answering
Fingerprint
Dive into the research topics of 'Looking closer and smarter: Multi-scale progressive attention for visual text question answering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver