Abstract
Large Vision Language Models (VLMs) are shown to be vulnerable to jailbreak attacks. Current attack detection methods often fall short due to their inability to comprehensively monitor the large input-output space or to accurately monitor safety-critical neurons associated with harmful semantics, resulting in high false positives and false negatives. In this paper, we introduce VLM-Guard, a highly effective detection framework that defends against jailbreak attacks by precisely identifying critical neurons linked to unsafe behaviors. Leveraging a tailored differential analysis over a large corpus of activation values, VLM-Guard isolates a compact set of neurons - just a few hundred, comprising less than 0.2% of the total - that are strongly correlated with harmful semantics. This enables the design of an attack detector that is not only effective at monitoring adversarial behavior but also lightweight and training-free (i.e., no parameter updates or model fine-tuning), making it well-suited for practical deployment. Extensive evaluations demonstrate that VLM-Guard excels in detecting jailbreak attacks while preserving benign performance in attack-free settings, offering an effective and efficient solution for safeguarding VLMs.
| Original language | English |
|---|---|
| Pages (from-to) | 4741-4754 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Information Forensics and Security |
| Volume | 21 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- Vision language models
- jailbreak attack and defense
- neuron activation
Fingerprint
Dive into the research topics of 'VLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver