Skip to main navigation Skip to search Skip to main content

Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought

  • Faculty of Computing, Harbin Institute of Technology
  • Qilu University of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

With the rapid development and widespread adoption of large language models (LLMs), the safety of LLMs has become a major concern. The inexplicability and unsafe outputs of LLMs pose significant obstacles to achieving artificial general intelligence (AGI). To enhance the safety of LLMs, researchers have developed various jailbreak attack methods and defense methods. In this paper, we propose SafeCoT, a novel defense method leveraging Chain-of-Thought (CoT) without any optimization or training. We believe that certain jailbreak attacks share a common logic, and based on this insight, we present SafeCoT. Specifically, to help LLMs understand the thoughts behind jailbreak attacks, we propose a jailbreak attack taxonomy and a corresponding jailbreak prompts dataset, JATD. Subsequently, we introduce SafeCoT, which consists of two parts: System Prompt and Safe Suffix. For different scenarios, we develop two forms of Safe Suffix, Manual-CoT and Zero-Shot-CoT. Through extensive experiments on 10 jailbreak attacks and 3 different LLMs, the results demonstrate that SafeCoT significantly reduces the attack success rate while maintaining good general performance. We hope our work can provide new perspectives and insights into LLM safety, and encourage further research to explore the underlying logic and mechanisms of jailbreak attacks.

Original languageEnglish
Title of host publicationNatural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
EditorsXian-Ling Mao, Zhaochun Ren, Muyun Yang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages196-207
Number of pages12
ISBN (Print)9789819533510
DOIs
StatePublished - 2026
Externally publishedYes
Event14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025 - Urumqi, China
Duration: 7 Aug 20259 Aug 2025

Publication series

NameLecture Notes in Computer Science
Volume16105 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
Country/TerritoryChina
CityUrumqi
Period7/08/259/08/25

Keywords

  • Chain-of-Thought
  • Jailbreak attack
  • LLM safety

Fingerprint

Dive into the research topics of 'Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought'. Together they form a unique fingerprint.

Cite this