TY - GEN
T1 - Thoughts Behind Attack
T2 - 14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
AU - Tao, Zhe
AU - Xu, Bing
AU - Yang, Muyun
AU - Guan, Hongjiao
AU - Lu, Wenpeng
AU - Cao, Hailong
AU - Zhu, Conghui
AU - Zhao, Tiejun
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd.
PY - 2026
Y1 - 2026
N2 - With the rapid development and widespread adoption of large language models (LLMs), the safety of LLMs has become a major concern. The inexplicability and unsafe outputs of LLMs pose significant obstacles to achieving artificial general intelligence (AGI). To enhance the safety of LLMs, researchers have developed various jailbreak attack methods and defense methods. In this paper, we propose SafeCoT, a novel defense method leveraging Chain-of-Thought (CoT) without any optimization or training. We believe that certain jailbreak attacks share a common logic, and based on this insight, we present SafeCoT. Specifically, to help LLMs understand the thoughts behind jailbreak attacks, we propose a jailbreak attack taxonomy and a corresponding jailbreak prompts dataset, JATD. Subsequently, we introduce SafeCoT, which consists of two parts: System Prompt and Safe Suffix. For different scenarios, we develop two forms of Safe Suffix, Manual-CoT and Zero-Shot-CoT. Through extensive experiments on 10 jailbreak attacks and 3 different LLMs, the results demonstrate that SafeCoT significantly reduces the attack success rate while maintaining good general performance. We hope our work can provide new perspectives and insights into LLM safety, and encourage further research to explore the underlying logic and mechanisms of jailbreak attacks.
AB - With the rapid development and widespread adoption of large language models (LLMs), the safety of LLMs has become a major concern. The inexplicability and unsafe outputs of LLMs pose significant obstacles to achieving artificial general intelligence (AGI). To enhance the safety of LLMs, researchers have developed various jailbreak attack methods and defense methods. In this paper, we propose SafeCoT, a novel defense method leveraging Chain-of-Thought (CoT) without any optimization or training. We believe that certain jailbreak attacks share a common logic, and based on this insight, we present SafeCoT. Specifically, to help LLMs understand the thoughts behind jailbreak attacks, we propose a jailbreak attack taxonomy and a corresponding jailbreak prompts dataset, JATD. Subsequently, we introduce SafeCoT, which consists of two parts: System Prompt and Safe Suffix. For different scenarios, we develop two forms of Safe Suffix, Manual-CoT and Zero-Shot-CoT. Through extensive experiments on 10 jailbreak attacks and 3 different LLMs, the results demonstrate that SafeCoT significantly reduces the attack success rate while maintaining good general performance. We hope our work can provide new perspectives and insights into LLM safety, and encourage further research to explore the underlying logic and mechanisms of jailbreak attacks.
KW - Chain-of-Thought
KW - Jailbreak attack
KW - LLM safety
UR - https://www.scopus.com/pages/publications/105046003130
U2 - 10.1007/978-981-95-3352-7_16
DO - 10.1007/978-981-95-3352-7_16
M3 - 会议稿件
AN - SCOPUS:105046003130
SN - 9789819533510
T3 - Lecture Notes in Computer Science
SP - 196
EP - 207
BT - Natural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
A2 - Mao, Xian-Ling
A2 - Ren, Zhaochun
A2 - Yang, Muyun
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 7 August 2025 through 9 August 2025
ER -