TY - GEN
T1 - SHARP
T2 - 30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
AU - Wu, Junfei
AU - Yue, Ding
AU - Liu, Guofan
AU - Xia, Tianze
AU - Huang, Ziyue
AU - Sui, Dianbo
AU - Liu, Qiang
AU - Wu, Shu
AU - Wang, Liang
AU - Tan, Tieniu
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Despite their impressive capabilities, Large Vision-Language Models (LVLMs) frequently generate plausible yet incorrect or unsupported responses, referred to as hallucinations. In this study, we investigate whether different types of hallucinations are reflected in the model's internal representations by probing their encoded features. We focus on two causes of hallucination in multimodal reasoning-(1) over-reliance on textual priors and (2) preference for user prompts over conflicting visual evidence-which have been identified in prior work as frequent and impactful factors. Our probing results reveals that hallucinations exhibit distinguishable representational patterns, suggesting a representation-level approach to characterize and mitigate them. Motivated by this, we propose Steering HAllucination via RePresentation Engineering (SHARP), a representation-level intervention framework that modulates hallucination-related features during inference. SHARP identifies functional representations responsible for prior-driven and visual-context conflicts, and jointly adjusts the model's internal activations during inference. We evaluate our approach extensively using three large vision-language models across various benchmarks. Experimental results show that our proposed intervention effectively reduces hallucinations without compromising the performance and generalization of the LVLMs.
AB - Despite their impressive capabilities, Large Vision-Language Models (LVLMs) frequently generate plausible yet incorrect or unsupported responses, referred to as hallucinations. In this study, we investigate whether different types of hallucinations are reflected in the model's internal representations by probing their encoded features. We focus on two causes of hallucination in multimodal reasoning-(1) over-reliance on textual priors and (2) preference for user prompts over conflicting visual evidence-which have been identified in prior work as frequent and impactful factors. Our probing results reveals that hallucinations exhibit distinguishable representational patterns, suggesting a representation-level approach to characterize and mitigate them. Motivated by this, we propose Steering HAllucination via RePresentation Engineering (SHARP), a representation-level intervention framework that modulates hallucination-related features during inference. SHARP identifies functional representations responsible for prior-driven and visual-context conflicts, and jointly adjusts the model's internal activations during inference. We evaluate our approach extensively using three large vision-language models across various benchmarks. Experimental results show that our proposed intervention effectively reduces hallucinations without compromising the performance and generalization of the LVLMs.
UR - https://www.scopus.com/pages/publications/105040261774
U2 - 10.18653/v1/2025.emnlp-main.725
DO - 10.18653/v1/2025.emnlp-main.725
M3 - 会议稿件
AN - SCOPUS:105040261774
T3 - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
SP - 14346
EP - 14361
BT - EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
A2 - Christodoulopoulos, Christos
A2 - Chakraborty, Tanmoy
A2 - Rose, Carolyn
A2 - Peng, Violet
PB - Association for Computational Linguistics (ACL)
Y2 - 4 November 2025 through 9 November 2025
ER -