Skip to main content
QUICK REVIEW

[Paper Review] Evaluating Explanation Without Ground Truth in Interpretable Machine Learning

Fan Yang, Mengnan Du|arXiv (Cornell University)|Jul 16, 2019
Explainable Artificial Intelligence (XAI)Computer Science70 references38 citations
TL;DR

This paper defines and investigates how to evaluate explanations in interpretable ML without ground-truth explanations, proposing generalizable, faithful, and persuasive criteria and a unified hierarchical framework for evaluation.

ABSTRACT

Interpretable Machine Learning (IML) has become increasingly important in many real-world applications, such as autonomous cars and medical diagnosis, where explanations are significantly preferred to help people better understand how machine learning systems work and further enhance their trust towards systems. However, due to the diversified scenarios and subjective nature of explanations, we rarely have the ground truth for benchmark evaluation in IML on the quality of generated explanations. Having a sense of explanation quality not only matters for assessing system boundaries, but also helps to realize the true benefits to human users in practical settings. To benchmark the evaluation in IML, in this article, we rigorously define the problem of evaluating explanations, and systematically review the existing efforts from state-of-the-arts. Specifically, we summarize three general aspects of explanation (i.e., generalizability, fidelity and persuasibility) with formal definitions, and respectively review the representative methodologies for each of them under different tasks. Further, a unified evaluation framework is designed according to the hierarchical needs from developers and end-users, which could be easily adopted for different scenarios in practice. In the end, open problems are discussed, and several limitations of current evaluation techniques are raised for future explorations.

Motivation & Objective

  • Clarify the problem of evaluating explanations in IML without ground truth ground truths.
  • Define three core properties of explanations: generalizability, fidelity, and persuasibility.
  • Review existing evaluation methods across different explanation types and applications.
  • Propose a unified, hierarchical evaluation framework aligned with developer and end-user needs.

Proposed method

  • Classify explanations using a two-dimensional scheme: interpretation scope (global/local) and interpretation manner (intrinsic/posthoc).
  • Formally define generalizability, fidelity, and persuasibility with precise definitions.
  • Systematically review existing evaluation methodologies corresponding to each property across tasks.
  • Propose a unified hierarchical evaluation framework with three tiers corresponding to generalizability, fidelity, and persuasibility.
  • Discuss open problems, limitations, and future directions for benchmarking explanation evaluation.

Experimental results

Research questions

  • RQ1How can explanations in IML be evaluated when there is no ground-truth explanation?
  • RQ2What formal properties best capture explanation quality in IML across tasks?
  • RQ3How can a unified framework support benchmarking explanations for both developers and end-users?
  • RQ4What are the key open problems and limitations in current evaluation techniques for explanations?
  • RQ5How should evaluation frameworks handle local vs. global and intrinsic vs. posthoc explanations?

Key findings

  • Three general properties—generalizability, fidelity, and persuasibility—are defined as core criteria for evaluating IML explanations.
  • Generalizability can align with traditional model evaluation for intrinsic-global explanations and with surrogate proxies for posthoc-global explanations.
  • Fidelity measures the faithfulness of explanations to the target system, with ablation/perturbation methods used for posthoc-local explanations.
  • Persuasibility assesses human usefulness and comprehensibility, often requiring human studies or annotations.
  • A unified hierarchical framework is proposed to organize evaluation from bottom (generalizability) to top (persuasibility), tailored to developer vs. end-user needs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.