Abstract
While aligning large language models (LLMs) with human preferences can prevent undesired outputs, recent research indicates that jailbreak prompts can easily bypass such safeguards, resulting in the generation of prohibited content. To counteract attacks, existing methods typically optimize LLMs’ decoding process or manipulate their input or output for safe responses. However, the scalability of the former is limited by their complexity, often requiring repeated generation or extensive retraining. In contrast, while the latter methods adopt simpler designs, they struggle against increasingly sophisticated attacks. Furthermore, most methods fail to intercept malicious queries upfront, requiring response generation before evaluating harmfulness, which adds significant overhead. In response, we propose an efficient Chain-of-Detection (CoD) mechanism for robust defense against evolving jailbreak attacks. Specifically, our approach first employs query-side detection to intercept common jailbreak attacks, using guided harmful search to promptly identify and block malicious intent. For more complex attacks that evade query-side detection, we leverage LLMs for deep reasoning on potentially harmful responses, enabling comprehensive defense by uncovering harmful queries at the response side. Experimental validation on datasets featuring diverse attacks across multiple closed-source and open-source LLMs confirms the effectiveness of our approach.
| Original language | English |
|---|---|
| Article number | 108217 |
| Journal | Neural Networks |
| Volume | 196 |
| DOIs | |
| State | Published - Apr 2026 |
Keywords
- Efficiency
- Jailbreak attacks
- Jailbreak defenses
- Large language models (LLMs)
- Robustness
Fingerprint
Dive into the research topics of 'Chain-of-Detection enables robust and efficient jailbreak defense'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver