Skip to main navigation Skip to search Skip to main content

Exploiting Shared Adversarial Features for Dynamic Attacks in Large Vision-Language Models

  • Yaguan Qian
  • , Xucheng Zhu
  • , Qiqi Bao*
  • , Fei Yu
  • , Shouling Ji
  • , Zhaoquan Gu*
  • , Wei Wang
  • , Bin Wang
  • , Zhen Lei
  • *Corresponding author for this work
  • Zhejiang University of Science and Technology
  • Liaoning University of Technology
  • Zhejiang University
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Xi'an Jiaotong University
  • Zhejiang Key Laboratory of Artificial Intelligence of Things (AIoT) Network and Data Security
  • CAS - Institute of Automation

Research output: Contribution to journalArticlepeer-review

Abstract

With the rapid development of Large Language Models (LLMs), an increasing number of Large Visual-Language Models (LVLMs) have achieved unprecedented performance in response generation. Recent work shows that LVLMs are vulnerable to adversarial attacks. However, many existing methods tend to overfit to the source model by overemphasizing specific features, which compromises their transferability. Other approaches suffer from reduced attack effectiveness due to insufficient differentiation between features. In this paper, we propose a novel transfer-based black-box untargeted attack—Shared Adversarial Feature (SAF) dynamic attack. By exploring the feature extraction patterns of LVLMs, we identify the features shared among various models that are most susceptible to adversarial attacks and disrupt them. Moreover, due to the powerful attention mechanisms of LVLMs, they are still able to extract similar semantics from perturbed images, even when primary features are disrupted. We design a dynamic update strategy to address this challenge. Finally, from the perspective of SAF, we conduct an in-depth analysis of vulnerabilities in the vision encoder and projector within LVLMs and find that attacking the projector exhibits stronger transferability across heterogeneous model architectures. Extensive experiments show that our method exhibits superior attack performance compared to existing methods across different models, datasets, and tasks.

Original languageEnglish
Pages (from-to)592-607
Number of pages16
JournalIEEE Transactions on Information Forensics and Security
Volume21
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Adversarial example
  • black-box attack
  • large visual-language model
  • model robustness
  • model security

Fingerprint

Dive into the research topics of 'Exploiting Shared Adversarial Features for Dynamic Attacks in Large Vision-Language Models'. Together they form a unique fingerprint.

Cite this