Skip to main navigation Skip to search Skip to main content

BAARD: Blocking Adversarial Examples by Testing for Applicability, Reliability and Decidability

  • Xinglong Chang*
  • , Katharina Dost
  • , Kaiqi Zhao
  • , Ambra Demontis
  • , Fabio Roli
  • , Gillian Dobbie
  • , Jörg Wicker
  • *Corresponding author for this work
  • The University of Auckland
  • University of Cagliari
  • University of Genoa

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Adversarial defenses protect machine learning models from adversarial attacks, but are often tailored to one type of model or attack. The lack of information on unknown potential attacks makes detecting adversarial examples challenging. Additionally, attackers do not need to follow the rules made by the defender. To address this problem, we take inspiration from the concept of Applicability Domain in cheminformatics. Cheminformatics models struggle to make accurate predictions because only a limited number of compounds are known and available for training. Applicability Domain defines a domain based on the known compounds and rejects any unknown compound that falls outside the domain. Similarly, adversarial examples start as harmless inputs, but can be manipulated to evade reliable classification by moving outside the domain of the classifier. We are the first to identify the similarity between Applicability Domain and adversarial detection. Instead of focusing on unknown attacks, we focus on what is known, the training data. We propose a simple yet robust triple-stage data-driven framework that checks the input globally and locally, and confirms that they are coherent with the model’s output. This framework can be applied to any classification model and is not limited to specific attacks. We demonstrate these three stages work as one unit, effectively detecting various attacks, even for a white-box scenario.

Original languageEnglish
Title of host publicationAdvances in Knowledge Discovery and Data Mining - 27th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2023, Proceedings
EditorsHisashi Kashima, Tsuyoshi Ide, Wen-Chih Peng
PublisherSpringer Science and Business Media Deutschland GmbH
Pages3-14
Number of pages12
ISBN (Print)9783031333736
DOIs
StatePublished - 2023
Externally publishedYes
Event27th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2023 - Hybrid, Osaka, Japan
Duration: 25 May 202328 May 2023

Publication series

NameLecture Notes in Computer Science
Volume13935 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference27th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2023
Country/TerritoryJapan
CityHybrid, Osaka
Period25/05/2328/05/23

Keywords

  • Adversarial Defense
  • Anomaly Detection
  • Applicability Domain
  • Evasion Attacks
  • White-box Adaptive Attacks

Fingerprint

Dive into the research topics of 'BAARD: Blocking Adversarial Examples by Testing for Applicability, Reliability and Decidability'. Together they form a unique fingerprint.

Cite this