arrow
Return

Considerations when learning additive explanations for black-box models

delete2023-06-19
delete6
PRE
AI
S
Sarah Tan *
G
Giles Hooker
P
Paul Koch
A
Albert Gordo
R
Rich Caruana
DOI:10.1007/s10994-023-06335-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Many methods to explain black-box models, whether local or global, are additive. In this paper, we study global additive explanations for non-additive models, focusing on four explanation methods: partial dependence, Shapley explanations adapted to a global setting, distilled additive explanations, and gradient-based explanations. We show that different explanation methods characterize non-additive components in a black-box model's prediction function in different ways. We use the concepts of main and total effects to anchor additive explanations, and quantitatively evaluate additive and non-additive explanations. Even though distilled explanations are generally the most accurate additive explanations, non-additive explanations such as tree explanations that explicitly model non-additive components tend to be even more accurate. Despite this, our user study showed that machine learning practitioners were better able to leverage additive explanations for various tasks. These considerations should be taken into account when considering which explanation to trust and use to explain black-box models.
Keywords:
Black-box models
Additive explanations
Model distillation
Interaction effects
Correlated features

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

F
facebook inc
Scholars:
588
Papers: 381
Citations: 0
U
University of California Berkeley
Scholars:
3.5W
Papers: 2.8W
Citations: 11.3W
University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K
researcher View more organizations