Skip to main navigation Skip to search Skip to main content

Robustness Evaluation of ChatGPT Against Chinese Adversarial Attacks

  • Yun Ting ZHANG
  • , Lin YE*
  • , Bai Song LI
  • , Hong Li ZHANG
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Antiy Labs

Research output: Contribution to journalArticlepeer-review

Abstract

Large language model (LLM) like ChatGPT has found widespread applications across various fields due to their strong natural language understanding and generation capabilities. However, deep learning models exhibit vulnerability when subjected to adversarial example attacks. In natural language processing, current research on adversarial example generation methods typically employs CNN-based models, RNN-based models, and Transformer-based pre-trained models as target models, with few studies exploring the robustness of LLMs under adversarial attacks and quantifying the evaluation criteria of LLM robustness. Taking ChatGPT against Chinese adversarial attacks as an example, this study introduces a novel concept termed offset average difference (OAD) and proposes a quantifiable LLM robustness evaluation metric based on OAD, named OAD-based robustness score (ORS). In a black-box attack scenario, this study selects nine mainstream Chinese adversarial attack methods based on word importance to generate adversarial texts, which are then employed to attack ChatGPT and yield the attack success rate of each method. The proposed ORS assigns a robustness score to LLMs for each attack method based on the attack success rate. In addition to the ChatGPT that outputs hard labels, this study designs ORS for target models with soft-labeled outputs based on the attack success rate and the proportion of misclassified adversarial texts with high confidence. Meanwhile, this study extends the scoring formula to the fluency assessment of adversarial texts, proposing an OAD-based adversarial text fluency scoring method, named OAD-based fluency score (OFS). Compared to traditional methods requiring human involvement, the proposed OFS greatly reduces evaluation costs. Experiments conducted on real-world Chinese news and sentiment classification datasets to some extent initially demonstrate that, for text classification tasks, the robustness score of ChatGPT against adversarial attacks is nearly 20% higher than that of Chinese BERT. However, the powerful ChatGPT still produces erroneous predictions under adversarial attacks, with the highest attack success rate exceeding 40%.

Translated title of the contribution中文对抗攻击下的 ChatGPT 鲁棒性评估
Original languageEnglish
Pages (from-to)4710-4734
Number of pages25
JournalRuan Jian Xue Bao/Journal of Software
Volume36
Issue number10
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • adversarial example (AE)
  • ChatGPT
  • ChatGPT
  • deep neural network (DNN)
  • large language model (LLM)
  • robustness
  • 大语言模型
  • 对抗样本
  • 深度神经网络
  • 鲁棒性

Fingerprint

Dive into the research topics of 'Robustness Evaluation of ChatGPT Against Chinese Adversarial Attacks'. Together they form a unique fingerprint.

Cite this