TY - GEN
T1 - Attack-words Guided Sentence Generation for Textual Adversarial Attack
AU - Zhang, Huan
AU - Xie, Yushun
AU - Zhu, Ziqi
AU - Sun, Jingling
AU - Li, Chao
AU - Gu, Zhaoquan
N1 - Publisher Copyright:
© 2021 IEEE.
PY - 2021
Y1 - 2021
N2 - Deep neural networks are vulnerable to carefully crafted adversarial examples and many adversarial attack methods have been proposed in computer vision tasks, such as image classification, object detection, etc. Generating adversarial examples for textual tasks is more challenging since the lexical correctness, grammatical correctness and semantics similarity should be maintained. In this paper, we introduce an attack-words guided sentence generation (AGSG) method to attack text classification models. We first determine words' attack ability by the ensemble strategy, then we add perturbation by inserting a short attack sentence. We conduct extensive experiments on two popular datasets IMDB and Amazon Comments against TextCNN, LSTM and RCNN models. The results show that the AGSG method greatly reduces the classification accuracy with a low word substitution rate. Specifically, the accuracy is reduced by 94.5% and 90.1% when disturbance rate is 13.3% and 25.1% for IMDB and Amazon Comments respectively. The similarity evaluation study shows that our adversarial attack method guarantees semantic similarity and grammatical correctness. Compared with two baseline adversarial attack methods, the AGSG method can generate adversarial texts that are harder for humans to perceive.
AB - Deep neural networks are vulnerable to carefully crafted adversarial examples and many adversarial attack methods have been proposed in computer vision tasks, such as image classification, object detection, etc. Generating adversarial examples for textual tasks is more challenging since the lexical correctness, grammatical correctness and semantics similarity should be maintained. In this paper, we introduce an attack-words guided sentence generation (AGSG) method to attack text classification models. We first determine words' attack ability by the ensemble strategy, then we add perturbation by inserting a short attack sentence. We conduct extensive experiments on two popular datasets IMDB and Amazon Comments against TextCNN, LSTM and RCNN models. The results show that the AGSG method greatly reduces the classification accuracy with a low word substitution rate. Specifically, the accuracy is reduced by 94.5% and 90.1% when disturbance rate is 13.3% and 25.1% for IMDB and Amazon Comments respectively. The similarity evaluation study shows that our adversarial attack method guarantees semantic similarity and grammatical correctness. Compared with two baseline adversarial attack methods, the AGSG method can generate adversarial texts that are harder for humans to perceive.
KW - Adversarial examples
KW - deep learning
KW - sentence generation
KW - text categorization
UR - https://www.scopus.com/pages/publications/85128747714
U2 - 10.1109/DSC53577.2021.00045
DO - 10.1109/DSC53577.2021.00045
M3 - 会议稿件
AN - SCOPUS:85128747714
T3 - Proceedings - 2021 IEEE 6th International Conference on Data Science in Cyberspace, DSC 2021
SP - 280
EP - 287
BT - Proceedings - 2021 IEEE 6th International Conference on Data Science in Cyberspace, DSC 2021
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 6th IEEE International Conference on Data Science in Cyberspace, DSC 2021
Y2 - 9 October 2021 through 11 October 2021
ER -