Skip to main navigation Skip to search Skip to main content

VLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons

  • Jinyin Hu
  • , Jiawei Zhou
  • , Minshan Xie
  • , Zhonghao Yang
  • , Jing Li
  • , Huadi Zheng
  • , Jie Shi
  • , Daojing He*
  • , Yu Li*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Guangdong University of Technology
  • East China Normal University
  • Huawei Technologies Co., Ltd.
  • Zhejiang University

Research output: Contribution to journalArticlepeer-review

Abstract

Large Vision Language Models (VLMs) are shown to be vulnerable to jailbreak attacks. Current attack detection methods often fall short due to their inability to comprehensively monitor the large input-output space or to accurately monitor safety-critical neurons associated with harmful semantics, resulting in high false positives and false negatives. In this paper, we introduce VLM-Guard, a highly effective detection framework that defends against jailbreak attacks by precisely identifying critical neurons linked to unsafe behaviors. Leveraging a tailored differential analysis over a large corpus of activation values, VLM-Guard isolates a compact set of neurons - just a few hundred, comprising less than 0.2% of the total - that are strongly correlated with harmful semantics. This enables the design of an attack detector that is not only effective at monitoring adversarial behavior but also lightweight and training-free (i.e., no parameter updates or model fine-tuning), making it well-suited for practical deployment. Extensive evaluations demonstrate that VLM-Guard excels in detecting jailbreak attacks while preserving benign performance in attack-free settings, offering an effective and efficient solution for safeguarding VLMs.

Original languageEnglish
Pages (from-to)4741-4754
Number of pages14
JournalIEEE Transactions on Information Forensics and Security
Volume21
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Vision language models
  • jailbreak attack and defense
  • neuron activation

Fingerprint

Dive into the research topics of 'VLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons'. Together they form a unique fingerprint.

Cite this