Abstract
Adversarial example detection is an important defense for safeguarding deep neural networks. However, since many existing detectors are themselves DNN-based models, attackers can apply gradient-based optimization strategies similar to those used against classifiers to craft adversarial examples (AEs) that both fool the target classifier and evade detection, thereby introducing an additional attack surface into the defense system. To address this challenge, we propose an AE detection framework based on implicit keys, which introduces a set of external random keys as hidden reference points. Instead of relying only on the input representation itself, the detector measures the relative distance distribution between an input and the implicit keys in a contrastive embedding space. We observe that natural examples (NEs) tend to preserve a relatively uniform relationship with these keys, whereas AEs exhibit more uneven distance distributions with respect to the keys. Leveraging this insight, we define topological distortion by characterizing the unevenness of the distance distribution between inputs and implicit keys, and train a dual-encoder detector to minimize the distortion of NEs while maximizing that of AEs. Since the implicit keys are independent of the training data and hidden from attackers, the proposed design reduces the exposed attack surface and improves robustness against adaptive attacks. Extensive experiments on CIFAR-10, CIFAR-100, SVHN, and ImageNet-100 under eight adversarial attack settings show that our method achieves an average AUROC of 98.63% and ranks first in 23 out of 32 evaluation settings. Moreover, the average AUROC degradation under adaptive attacks is only 0.5% in the gray-box setting and 2.1% in the white-box setting, demonstrating the effectiveness and robustness of the proposed detector.
| Original language | English |
|---|---|
| Article number | 115890 |
| Journal | Applied Soft Computing |
| Volume | 202 |
| DOIs | |
| State | Published - Oct 2026 |
| Externally published | Yes |
Keywords
- Adversarial example detection
- Attack surface
- Contrastive learning
- Implicit keys
- Topological distortion
Fingerprint
Dive into the research topics of 'Reducing the attack surface of adversarial example detectors with implicit keys'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver