Skip to main content
QUICK REVIEW

[Paper Review] Metrics for Explainable AI: Challenges and Prospects

Robert R. Hoffman, Shane T. Mueller|arXiv (Cornell University)|Dec 11, 2018
Explainable Artificial Intelligence (XAI)Computer Science107 references220 citations
TL;DR

This paper investigates how to evaluate XAI by examining how to measure explanation quality, user satisfaction and understanding, curiosity-driven explanation seeking, appropriate trust and reliance, and overall human-XAI system performance, drawing on psychometrics and literature integration.

ABSTRACT

The question addressed in this paper is: If we present to a user an AI system that explains how it works, how do we know whether the explanation works and the user has achieved a pragmatic understanding of the AI? In other words, how do we know that an explanainable AI system (XAI) is any good? Our focus is on the key concepts of measurement. We discuss specific methods for evaluating: (1) the goodness of explanations, (2) whether users are satisfied by explanations, (3) how well users understand the AI systems, (4) how curiosity motivates the search for explanations, (5) whether the user's trust and reliance on the AI are appropriate, and finally, (6) how the human-XAI work system performs. The recommendations we present derive from our integration of extensive research literatures and our own psychometric evaluations.

Motivation & Objective

  • Motivate the need for measurement in explainable AI (XAI) to ensure pragmatic user understanding.
  • Identify key measurement targets for XAI evaluations (explanation goodness, user satisfaction, understanding, curiosity, trust/reliance, system performance).
  • Synthesize insights from extensive literatures and psychometric work to guide evaluation practices.

Proposed method

  • Integrates diverse research literatures on XAI and measurement.
  • Proposes evaluation concepts and domains grounded in psychometric evaluation and user studies.
  • Offers recommendations for assessing multiple facets of the user–AI explanatory loop.

Experimental results

Research questions

  • RQ1How can we assess the goodness of explanations provided by AI systems?
  • RQ2How well do users feel satisfied with explanations, and how does that relate to understanding?
  • RQ3To what extent do explanations stimulate curiosity and search for further information?
  • RQ4How appropriate are users' trust and reliance on the AI system given the explanations?
  • RQ5How should the human–XAI work system be evaluated as a whole?

Key findings

  • Provide a set of measurement targets spanning explanation quality, user satisfaction, understanding, curiosity, trust/reliance, and system-level performance.
  • Advocate for integrating extensive literature and psychometric evaluations to derive practical recommendations for XAI measurement.
  • Offer concrete recommendations for evaluating XAI effectiveness based on synthesized evidence and evaluation paradigms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.