[Paper Review] On the (In)fidelity and Sensitivity for Explanations
The paper formalizes an infidelity objective for saliency explanations of black-box models, derives optimal explanations under different perturbations, and shows smoothing-based approaches can reduce both sensitivity and infidelity, validated by experiments.
We consider objective evaluation measures of saliency explanations for complex black-box machine learning models. We propose simple robust variants of two notions that have been considered in recent literature: (in)fidelity, and sensitivity. We analyze optimal explanations with respect to both these measures, and while the optimal explanation for sensitivity is a vacuous constant explanation, the optimal explanation for infidelity is a novel combination of two popular explanation methods. By varying the perturbation distribution that defines infidelity, we obtain novel explanations by optimizing infidelity, which we show to out-perform existing explanations in both quantitative and qualitative measurements. Another salient question given these measures is how to modify any given explanation to have better values with respect to these measures. We propose a simple modification based on lowering sensitivity, and moreover show that when done appropriately, we could simultaneously improve both sensitivity as well as fidelity.
Motivation & Objective
- Motivate objective evaluation of saliency explanations for black-box models.
- Define and analyze a robust infidelity measure to quantify how explanations capture predictor changes under significant input perturbations.
- Relate infidelity to existing explanations and derive new perturbation-based explanations.
- Propose smoothing-based modifications to explanations that reduce both sensitivity and infidelity, with practical validation.
Proposed method
- Define explanation infidelity as the expected squared difference between perturbation-weighted explanation and actual function change under perturbations I.
- Characterize optimal explanations Φ* that minimize infidelity via an integral with perturbation distribution μI and Integrated Gradients (IG).
- Show that many existing explanations (IG, DeepLIFT, LRP) are optimal for infidelity under specific perturbations; derive novel explanations for other perturbations (e.g., noisy baseline, square removal).
- Propose kernel smoothing (Φk) to obtain smoother explanations, relate to Smooth-Grad, and provide conditions under which infidelity improves after smoothing.
- Introduce a robust, Monte Carlo-friendly max-sensitivity measure and relate it to fidelity via smoothing, including adversarial training as an optional enhancement.
Experimental results
Research questions
- RQ1What objective measures can quantify the faithfulness of saliency explanations to a black-box predictor?
- RQ2How do different perturbation schemes affect the optimal explanation under the infidelity objective, and can we design novel explanations from these perturbations?
- RQ3Can simple smoothing or training strategies reduce both the sensitivity and the infidelity of explanations without sacrificing fidelity?
- RQ4How do existing explanations perform under the infidelity framework, and do improvements in smoothing correlate with human judgments?
Key findings
- The optimal infidelity-minimizing explanation can be expressed as a smoothed Integrated Gradients-style combination using a perturbation-induced kernel.
- Many existing explanations (IG, DeepLIFT, LRP) arise as special cases of infidelity-optimal explanations under particular perturbations; new perturbations yield novel explanations.
- Smoothing-based adjustments (e.g., Smooth-Grad) reduce both sensitivity and infidelity in most cases, and can improve qualitative saliency maps.
- Relaxed, robust perturbations (e.g., noisy baseline, square removal) lead to explanations with lower infidelity and more faithful visualizations, confirmed by human evaluation.
- Adversarial training can also help lower both sensitivity and infidelity, suggesting model-level strategies for more faithful explanations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.