[Paper Review] Explanations can be manipulated and geometry is to blame
The paper argues that explanations produced by common attribution methods can be manipulated, and attributes this vulnerability to geometric properties of models and inputs.
Explanation methods aim to make neural networks more trustworthy and interpretable. In this paper, we demonstrate a property of explanation methods which is disconcerting for both of these purposes. Namely, we show that explanations can be manipulated arbitrarily by applying visually hardly perceptible perturbations to the input that keep the network's output approximately constant. We establish theoretically that this phenomenon can be related to certain geometrical properties of neural networks. This allows us to derive an upper bound on the susceptibility of explanations to manipulations. Based on this result, we propose effective mechanisms to enhance the robustness of explanations.
Motivation & Objective
- Motivate and analyze the susceptibility of explanation methods to manipulation.
- Examine how geometric properties of models and input spaces contribute to explanation vulnerabilities.
- Survey several standard explanation methods and discuss their limitations in the context of manipulation.
Proposed method
- Describe gradient-based attribution methods such as Gradient, Gradient × Input, and Integrated Gradients.
- Discuss backpropagation-based explanations including Guided Backpropagation and Layer-wise Relevance Propagation.
- Highlight how geometry of the input space and model decision boundaries influence explanation behavior.
Experimental results
Research questions
- RQ1Can common explanation methods be manipulated or deceived by adversarial inputs?
- RQ2What role does the geometry of the model and input space play in the reliability of explanations?
- RQ3Do popular attribution techniques have inherent vulnerabilities that enable manipulation?
- RQ4How do different attribution methods compare in their susceptibility to manipulation?
Key findings
- Explanations produced by attribution methods can be susceptible to manipulation.
- Geometry plays a central role in the vulnerability of explanations across methods.
- Several standard attribution techniques (e.g., Gradient, Gradient × Input, Integrated Gradients, GBP, LRP) are discussed in the context of their weaknesses.
- The paper analyzes how perturbations in pixels can affect the resulting explanations.
- The study connects the mathematical properties of backpropagation and relevance propagation to explainability weaknesses.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.