Skip to main navigation Skip to search Skip to main content

TF-Attack: Transferable and fast adversarial attacks on large language models

  • School of Computer Science and Technology, Harbin Institute of Technology
  • Independent
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

With the great advancements in large language models (LLMs), adversarial attacks against LLMs have recently attracted increasing attention. We found that pre-existing adversarial attack methodologies exhibit limited transferability and are notably inefficient, particularly when applied to LLMs. In this paper, we analyze the core mechanisms of previous predominant adversarial attack methods, revealing that (1) the distributions of importance score differ markedly among victim models, restricting the transferability; (2) the sequential attack processes induces substantial time overheads. Based on the above two insights, we introduce a new scheme, named TF-ATTACK, for Transferable and Fast adversarial attacks on LLMs. TF-ATTACK employs an external LLM as a third-party overseer rather than the victim model to identify critical units within sentences. Moreover, TF-ATTACK introduces the concept of Importance Level, which allows for parallel substitutions of attacks. We conduct extensive experiments on 6 widely adopted benchmarks, evaluating the proposed method through both automatic and human metrics. Results show that our method consistently surpasses previous methods in transferability and delivers significant speed improvements, up to 10× faster than earlier attack strategies.

Original languageEnglish
Article number113117
JournalKnowledge-Based Systems
Volume312
DOIs
StatePublished - 15 Mar 2025
Externally publishedYes

Keywords

  • Large language models
  • Natural language processing
  • Sample transferability
  • Text adversarial attacks

Fingerprint

Dive into the research topics of 'TF-Attack: Transferable and fast adversarial attacks on large language models'. Together they form a unique fingerprint.

Cite this