Skip to main content
QUICK REVIEW

[Paper Review] Towards A Rigorous Science of Interpretable Machine Learning

Finale Doshi‐Velez, Been Kim|arXiv (Cornell University)|Feb 28, 2017
Explainable Artificial Intelligence (XAI)26 references3,111 citations
TL;DR

Proposes a formal framework and taxonomy for evaluating interpretability in ML, linking applications, human studies, and proxy metrics, and outlining open problems and a research agenda.

ABSTRACT

As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other criteria such as safety or non-discrimination. However, despite the interest in interpretability, there is very little consensus on what interpretable machine learning is and how it should be measured. In this position paper, we first define interpretability and describe when interpretability is needed (and when it is not). Next, we suggest a taxonomy for rigorous evaluation and expose open questions towards a more rigorous science of interpretable machine learning.

Motivation & Objective

  • Define interpretability in ML and distinguish it from related criteria like reliability and fairness.
  • Argue the need for rigorous, evidence-based evaluation of interpretability.
  • Propose a taxonomy for evaluating interpretability: application-grounded, human-grounded, and functionally-grounded.
  • Outline open problems and data-driven approaches to uncover latent dimensions of interpretability.
  • Provide recommendations for researchers on how to report and frame interpretability work.

Proposed method

  • Define interpretability as the ability to explain or present in understandable terms to a human.
  • Introduce a three-tier taxonomy of evaluation: application-grounded, human-grounded, and functionally-grounded.
  • Discuss trade-offs and design considerations for human-subject experiments in interpretability.
  • Propose data-driven approaches to discover latent factors of interpretability, including task-method matrices and matrix factorization ideas.
  • Suggest hypotheses about task-related and method-related latent dimensions of interpretability.
  • Outline best practices for matching claims to appropriate evaluation types.

Experimental results

Research questions

  • RQ1What constitutes a rigorous, evidence-based evaluation of interpretability in ML?
  • RQ2How should interpretability be categorized to align evaluation with claims (application-specific vs general)?
  • RQ3What proxies or factors best capture interpretability across tasks and methods?
  • RQ4How can we link application-grounded, human-grounded, and functionally-grounded evaluations?
  • RQ5What open problems must be addressed to build a shared language and repositories for interpretability research?

Key findings

  • Interpretability lacks a single, universally agreed definition and requires formalization to enable meaningful comparison.
  • A taxonomy of evaluation approaches (application-grounded, human-grounded, functionally-grounded) is proposed to align evaluation with the type of claim.
  • Human evaluation is essential but challenging; different evaluation types incur different costs and biases.
  • Data-driven approaches (e.g., task-method matrices and embeddings) could uncover latent dimensions of interpretability and guide method selection.
  • Three open problems are identified: choosing appropriate proxies, designing simpler tasks that preserve end-task essence, and characterizing proxies for explanation quality.
  • The paper provides practical recommendations to ground interpretability work in a common taxonomy and avoid vague claims.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.