Skip to main navigation Skip to search Skip to main content

Looking closer and smarter: Multi-scale progressive attention for visual text question answering

  • Faculty of Computing, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Vision-language models have advanced rapidly, yet they often rely on coarse, global image features and struggle with fine-grained cross-modal reasoning. The Visual Text Question Answering (VTQA) benchmark exemplifies this challenge: each instance pairs a natural image with a detailed text passage, requiring the alignment of specific entities across modalities for multi-hop reasoning. Existing approaches often employ frozen vision backbones that capture only global semantics, missing critical local details. To address this, we propose Multi-Scale Progressive Attention (MSPA). MSPA encodes images at multiple resolutions to obtain multi-scale features and applies a hierarchical top-k selection mechanism: at each scale, the model conditions on the textual query to identify the most relevant visual tokens, which are then projected to the next finer scale for refinement. This coarse-to-fine strategy concentrates computation on key regions, enabling detailed visual understanding without processing all high-resolution tiles. On the VTQA benchmark, MSPA establishes a new state-of-the-art, achieving 71.2% Exact Match (EM) on the test set (surpassing the previous best of 66.4%). These results demonstrate that integrating multi-scale visual features with progressive attention substantially enhances fine-grained visual-text reasoning, suggesting that coarse-to-fine processing can benefit other vision-language tasks demanding precise grounding.

Original languageEnglish
Article number134131
JournalNeurocomputing
Volume697
DOIs
StatePublished - 7 Oct 2026
Externally publishedYes

Keywords

  • Multi-scale attention
  • Vision-language models
  • Visual question answering

Fingerprint

Dive into the research topics of 'Looking closer and smarter: Multi-scale progressive attention for visual text question answering'. Together they form a unique fingerprint.

Cite this