[Paper Review] Problems with Shapley-value-based explanations as feature importance measures
The paper critiques Shapley-value-based feature explanations, arguing mathematical and human-centric shortcomings, and discusses interventional vs conditional value functions, additivity limits, and implications for explainability.
Game-theoretic formulations of feature importance have become popular as a way to "explain" machine learning models. These methods define a cooperative game between the features of a model and distribute influence among these input elements using some form of the game's unique Shapley values. Justification for these methods rests on two pillars: their desirable mathematical properties, and their applicability to specific motivations for explanations. We show that mathematical problems arise when Shapley values are used for feature importance and that the solutions to mitigate these necessarily induce further complexity, such as the need for causal reasoning. We also draw on additional literature to argue that Shapley values do not provide explanations which suit human-centric goals of explainability.
Motivation & Objective
- Assess whether Shapley-value-based explanations reliably reflect feature importance for model explanations.
- Identify mathematical issues arising from using Shapley values for feature importance.
- Evaluate human-centric adequacy of Shapley-based explanations under established explainability frameworks.
- Suggest conditions or alternatives where Shapley-based explanations may be meaningful.
Proposed method
- Review and formal analysis of Shapley value foundations and their application to feature importance.
- Compare interventional versus conditional value functions for v_f,x and their computational implications.
- Examine additivity and other axioms in non-additive models to assess interpretability.
- Discuss causal considerations and how prior knowledge affects attribution (asymmetric Shapley values).
- Analyze human-centric perspectives on explanations using contrastive explanations and normative evaluation frameworks.
Experimental results
Research questions
- RQ1Do Shapley-value-based explanations align with human notions of explainability and contrastive reasoning?
- RQ2What mathematical issues arise from using conditional vs interventional value functions for feature importance?
- RQ3How do additivity and other axioms constrain the interpretability of Shapley-based attributions in non-additive models?
- RQ4Can Shapley-based explanations meaningfully support actionable recourse or normative evaluation without misalignment?
- RQ5Under what constrained settings might Shapley-based explanations admit clear interpretation?
Key findings
- Shapley-value explanations can attribute influence to features with no interventional effect due to conditional value functions.
- Interventional methods require out-of-distribution model evaluation, leading to misleading explanations for in-distribution samples.
- Additivity axioms constrain attribution in additive models and may be uninformative for non-additive models.
- Redundant or highly correlated features can distort attributions depending on feature inclusion choices.
- Causal knowledge can mitigate some issues but introduces dependence on prior assumptions and may undermine universality of explanations.
- Human-centric critiques suggest Shapley explanations often fail to meet the contrastive and actionable expectations central to explanations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.