Abstract
Currently, interpretability methods focus more on less objective human-understandable semantics. To objectify and standardize interpretability research, in this study, we provide notions of interpretability based on approximation theory. We first define explainable models in terms of explicitness and then use completeness to define interpretability, thereby turning interpretability into the process of approximating black-box models with interpretable models. In particular, we think that the decision boundary of a classification model is equivalent to its interpretability. Next, we implement this approximation interpretation on multilayer perceptrons (MLPs) and then propose to use the MLP as a universal interpreter to explain other complex black-box models. Compared to the LIME method, which can only extract local linear features, our method is global and therefore termed as GIME. Extensive experiments demonstrate the effectiveness of our approaches.
| Original language | English |
|---|---|
| Article number | 4339 |
| Journal | Electronics (Switzerland) |
| Volume | 13 |
| Issue number | 22 |
| DOIs | |
| State | Published - Nov 2024 |
| Externally published | Yes |
Keywords
- decision boundary
- deep learning
- interpretability
- knowledge distillation
Fingerprint
Dive into the research topics of 'Interpretability as Approximation: Understanding Black-Box Models by Decision Boundary'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver