Skip to main content
QUICK REVIEW

[論文レビュー] A Meta Survey of Quality Evaluation Criteria in Explanation Methods

Helena Löfström, Karl Hammar|arXiv (Cornell University)|Mar 25, 2022
Explainable Artificial Intelligence (XAI)参考文献 26被引用数 4
ひとこと要約

本稿では、XAIにおける説明手法の比較的評価を可能にするために、'適切な信頼'を測定可能な結果指標として提案する。15件の文献レビューを分析することで、モデル、説明、ユーザーの3つの品質的側面にわたる4つのコア基準—性能、適切な信頼、説明満足度、忠実度—を同定し、評価の標準化と比較研究における主観性の克服を図る統一的モデルを提示する。

ABSTRACT

Explanation methods and their evaluation have become a significant issue in explainable artificial intelligence (XAI) due to the recent surge of opaque AI models in decision support systems (DSS). Since the most accurate AI models are opaque with low transparency and comprehensibility, explanations are essential for bias detection and control of uncertainty. There are a plethora of criteria to choose from when evaluating explanation method quality. However, since existing criteria focus on evaluating single explanation methods, it is not obvious how to compare the quality of different methods. This lack of consensus creates a critical shortage of rigour in the field, although little is written about comparative evaluations of explanation methods. In this paper, we have conducted a semi-systematic meta-survey over fifteen literature surveys covering the evaluation of explainability to identify existing criteria usable for comparative evaluations of explanation methods. The main contribution in the paper is the suggestion to use appropriate trust as a criterion to measure the outcome of the subjective evaluation criteria and consequently make comparative evaluations possible. We also present a model of explanation quality aspects. In the model, criteria with similar definitions are grouped and related to three identified aspects of quality; model, explanation, and user. We also notice four commonly accepted criteria (groups) in the literature, covering all aspects of explanation quality: Performance, appropriate trust, explanation satisfaction, and fidelity. We suggest the model be used as a chart for comparative evaluations to create more generalisable research in explanation quality.

研究の動機と目的

  • XAIにおける説明手法の評価において、合意形成と標準化が不足している問題に対処する。
  • 既存のサーベイから共通して受け入れられている評価基準を同定し、比較的評価を可能にする。
  • 主観的なユーザー評価の課題を克服するため、'適切な信頼'を客観的な結果指標として提案する。
  • モデル、説明、ユーザーの側面にわたる基準を統合した説明品質の構造的モデルを構築する。
  • XAI研究における説明手法の一般化可能で比較可能な評価のフレームワークを提供する。

提案手法

  • XAIの説明手法評価に関する15件の文献レビューを対象に、準系統的なメタサーベイを実施した。
  • サーベイから抽出された評価基準を、共通する定義と目的に基づいて一貫したカテゴリに分類・統合した。
  • 説明品質の3つのコア側面(モデル、説明、ユーザー)を特定し、それらに関連する基準を同定した。
  • '適切な信頼'を主観的ユーザー基準の測定可能な結果指標として提案し、客観的な比較を可能にした。
  • 3つの側面にわたる11の評価基準グループを統合した、高レベルの説明品質モデルを構築した。
  • 4つの基準—性能、適切な信頼、説明満足度、忠実度—が半数以上のサーベイで登場し、3つの品質側面すべてにまたがることを示すことで、モデルの有効性を検証した。

実験結果

リサーチクエスチョン

  • RQ1説明手法に関する文献レビューにおいて、最も一貫して使用されている評価基準は何か?
  • RQ2主観的なユーザー評価基準を、説明手法の評価に適した客観的・比較可能な指標にどのように変換できるか?
  • RQ3説明品質のコア側面とは何か。また、評価基準はそれらの側面とどのように関連しているか?
  • RQ4既存の基準から、比較的評価を支援する統一的な説明品質モデルを構築できるか?
  • RQ5性能、忠実度、説明満足度、適切な信頼が、説明手法の評価における基盤的基準としてどの程度の役割を果たすか?

主な発見

  • 性能、適切な信頼、説明満足度、忠実度という4つの基準が、調査された15件の文献レビューの半数以上で一貫して言及されており、それらの重要性について広範な合意があることを示している。
  • メタサーベイにより、11の異なる評価基準グループが同定され、これらはモデル、説明、ユーザーの3つの品質側面に整理された。
  • 基準'適切な信頼'は、主観的ユーザー評価基準の成功を客観的に測定できる重要な結果指標であると特定され、異なる手法間の比較を可能にする。
  • 研究では、既存の評価実践が人間を含むフィードバックループに大きく依存しており、研究間での再現性と比較可能性の課題を生じさせていることが判明した。
  • 提示された説明品質モデルは、多様な評価基準を統一的に整列させ、XAIにおけるより一般化可能な研究を支援する構造的フレームワークを提供する。
  • 多くの基準(例:信頼性、自信、確実性)が研究間で入れ替え可能または曖昧に使用されていることから、標準化された定義とベンチマークの必要性が強調された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。