Skip to main navigation Skip to search Skip to main content

SHARP: Steering Hallucination in LVLMs via Representation Engineering

  • Junfei Wu
  • , Ding Yue
  • , Guofan Liu
  • , Tianze Xia
  • , Ziyue Huang
  • , Dianbo Sui
  • , Qiang Liu*
  • , Shu Wu
  • , Liang Wang
  • , Tieniu Tan
  • *Corresponding author for this work
  • CAS - Institute of Automation
  • University of Chinese Academy of Sciences
  • Nanjing University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Despite their impressive capabilities, Large Vision-Language Models (LVLMs) frequently generate plausible yet incorrect or unsupported responses, referred to as hallucinations. In this study, we investigate whether different types of hallucinations are reflected in the model's internal representations by probing their encoded features. We focus on two causes of hallucination in multimodal reasoning-(1) over-reliance on textual priors and (2) preference for user prompts over conflicting visual evidence-which have been identified in prior work as frequent and impactful factors. Our probing results reveals that hallucinations exhibit distinguishable representational patterns, suggesting a representation-level approach to characterize and mitigate them. Motivated by this, we propose Steering HAllucination via RePresentation Engineering (SHARP), a representation-level intervention framework that modulates hallucination-related features during inference. SHARP identifies functional representations responsible for prior-driven and visual-context conflicts, and jointly adjusts the model's internal activations during inference. We evaluate our approach extensively using three large vision-language models across various benchmarks. Experimental results show that our proposed intervention effectively reduces hallucinations without compromising the performance and generalization of the LVLMs.

Original languageEnglish
Title of host publicationEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
EditorsChristos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
PublisherAssociation for Computational Linguistics (ACL)
Pages14346-14361
Number of pages16
ISBN (Electronic)9798891763326
DOIs
StatePublished - 2025
Event30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - Suzhou, China
Duration: 4 Nov 20259 Nov 2025

Publication series

NameEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference

Conference

Conference30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Country/TerritoryChina
CitySuzhou
Period4/11/259/11/25

Fingerprint

Dive into the research topics of 'SHARP: Steering Hallucination in LVLMs via Representation Engineering'. Together they form a unique fingerprint.

Cite this