TY - GEN
T1 - Fine-Grained Attention Enhancement for Mitigating Hallucinations in LVLMs
AU - Yang, Jidong
AU - Yao, Hongxun
AU - Chen, Xi
AU - Jiang, Shouxu
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Large vision-language models (LVLMs) have achieved impressive performance on tasks such as visual question answering and image captioning, yet they still suffer from hallucinations, producing descriptions that conflict with the visual input. Prior efforts to mitigate hallucinations often rely on data-centric strategies or specialized fine-tuning, which require large-scale annotations or costly retraining. More recent training-free methods adjust attention at the modality level but largely ignore fine-grained object–region grounding. In this work, we analyze hallucinations through the lens of cross-modal attention and show that, during decoding, LVLMs tend to collapse onto a few visually global tokens, while progressive encoding compresses visual evidence into global tokens that are easily overridden by language priors. To address this, we propose a training-free framework that (i) leverages CLIP patch embeddings and early-layer attention maps to reweight decoder cross-attention in an object-centric, fine-grained manner, and (ii) introduces an auxiliary decoding branch with masked global tokens for contrastive decoding, effectively reducing hallucinations without additional training.
AB - Large vision-language models (LVLMs) have achieved impressive performance on tasks such as visual question answering and image captioning, yet they still suffer from hallucinations, producing descriptions that conflict with the visual input. Prior efforts to mitigate hallucinations often rely on data-centric strategies or specialized fine-tuning, which require large-scale annotations or costly retraining. More recent training-free methods adjust attention at the modality level but largely ignore fine-grained object–region grounding. In this work, we analyze hallucinations through the lens of cross-modal attention and show that, during decoding, LVLMs tend to collapse onto a few visually global tokens, while progressive encoding compresses visual evidence into global tokens that are easily overridden by language priors. To address this, we propose a training-free framework that (i) leverages CLIP patch embeddings and early-layer attention maps to reweight decoder cross-attention in an object-centric, fine-grained manner, and (ii) introduces an auxiliary decoding branch with masked global tokens for contrastive decoding, effectively reducing hallucinations without additional training.
KW - Hallucination
KW - LVLMs
KW - Training-free
UR - https://www.scopus.com/pages/publications/105040595687
U2 - 10.1007/978-981-95-9493-1_9
DO - 10.1007/978-981-95-9493-1_9
M3 - 会议稿件
AN - SCOPUS:105040595687
SN - 9789819594924
T3 - Communications in Computer and Information Science
SP - 134
EP - 148
BT - Emotional Intelligence - 3rd CSIG Conference, CEI 2025, Proceedings
A2 - Liu, Honghai
A2 - Yao, Hongxun
A2 - Zhang, Shengping
A2 - Ren, Weihong
A2 - Wang, Zhiyong
A2 - Huang, Hui
PB - Springer Science and Business Media Deutschland GmbH
T2 - 3rd CSIG Conference on Emotional Intelligence, CEI 2025
Y2 - 5 December 2025 through 7 December 2025
ER -