Skip to main navigation Skip to search Skip to main content

Fine-Grained Attention Enhancement for Mitigating Hallucinations in LVLMs

  • Jidong Yang
  • , Hongxun Yao*
  • , Xi Chen
  • , Shouxu Jiang
  • *Corresponding author for this work
  • Faculty of Computing, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large vision-language models (LVLMs) have achieved impressive performance on tasks such as visual question answering and image captioning, yet they still suffer from hallucinations, producing descriptions that conflict with the visual input. Prior efforts to mitigate hallucinations often rely on data-centric strategies or specialized fine-tuning, which require large-scale annotations or costly retraining. More recent training-free methods adjust attention at the modality level but largely ignore fine-grained object–region grounding. In this work, we analyze hallucinations through the lens of cross-modal attention and show that, during decoding, LVLMs tend to collapse onto a few visually global tokens, while progressive encoding compresses visual evidence into global tokens that are easily overridden by language priors. To address this, we propose a training-free framework that (i) leverages CLIP patch embeddings and early-layer attention maps to reweight decoder cross-attention in an object-centric, fine-grained manner, and (ii) introduces an auxiliary decoding branch with masked global tokens for contrastive decoding, effectively reducing hallucinations without additional training.

Original languageEnglish
Title of host publicationEmotional Intelligence - 3rd CSIG Conference, CEI 2025, Proceedings
EditorsHonghai Liu, Hongxun Yao, Shengping Zhang, Weihong Ren, Zhiyong Wang, Hui Huang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages134-148
Number of pages15
ISBN (Print)9789819594924
DOIs
StatePublished - 2026
Externally publishedYes
Event3rd CSIG Conference on Emotional Intelligence, CEI 2025 - Shenzhen, China
Duration: 5 Dec 20257 Dec 2025

Publication series

NameCommunications in Computer and Information Science
Volume2881 CCIS
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference3rd CSIG Conference on Emotional Intelligence, CEI 2025
Country/TerritoryChina
CityShenzhen
Period5/12/257/12/25

Keywords

  • Hallucination
  • LVLMs
  • Training-free

Fingerprint

Dive into the research topics of 'Fine-Grained Attention Enhancement for Mitigating Hallucinations in LVLMs'. Together they form a unique fingerprint.

Cite this