Skip to main navigation Skip to search Skip to main content

Interpretability as Approximation: Understanding Black-Box Models by Decision Boundary

  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Currently, interpretability methods focus more on less objective human-understandable semantics. To objectify and standardize interpretability research, in this study, we provide notions of interpretability based on approximation theory. We first define explainable models in terms of explicitness and then use completeness to define interpretability, thereby turning interpretability into the process of approximating black-box models with interpretable models. In particular, we think that the decision boundary of a classification model is equivalent to its interpretability. Next, we implement this approximation interpretation on multilayer perceptrons (MLPs) and then propose to use the MLP as a universal interpreter to explain other complex black-box models. Compared to the LIME method, which can only extract local linear features, our method is global and therefore termed as GIME. Extensive experiments demonstrate the effectiveness of our approaches.

Original languageEnglish
Article number4339
JournalElectronics (Switzerland)
Volume13
Issue number22
DOIs
StatePublished - Nov 2024
Externally publishedYes

Keywords

  • decision boundary
  • deep learning
  • interpretability
  • knowledge distillation

Fingerprint

Dive into the research topics of 'Interpretability as Approximation: Understanding Black-Box Models by Decision Boundary'. Together they form a unique fingerprint.

Cite this