Abstract
Vision Transformers (ViTs) have achieved state-of-the-art performance across a wide range of vision tasks, yet they remain highly vulnerable to adversarial patch attacks—deliberate image modifications confined to small, localized regions (patches) that can completely mislead the model. Existing defenses often treat all adversarial inputs uniformly, which either removes essential features or fails to eliminate malicious influence, thereby limiting robustness. We present AdaGuard, an adaptive defense framework that combines adversarial patch detection with location-aware treatment. AdaGuard first employs a lightweight, model-agnostic detector based on patch–neighbor discontinuity to accurately localize adversarial regions of unknown size and position. It then distinguishes between non-catastrophic attacks (patches in less informative regions) and catastrophic attacks (patches overlapping with essential features). For non-catastrophic cases, AdaGuard reconstructs the region with neigh-boring content to fully remove adversarial information, while for catastrophic cases it applies fine-grained attention suppression to neutralize adversarial influence without discarding critical features. Experiments on 12 ViT models and multiple attack methods on ImageNet demonstrate that AdaGuard achieves state-of-the-art robust accuracy while preserving clean accuracy, offering a practical defense for ViTs in real-world applications.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- adaptive defense
- adversarial patch attack
- vision transformers
Fingerprint
Dive into the research topics of 'AdaGuard: Location-Aware Adaptive Defense against Adversarial Patch Attacks in Vision Transformers'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver