[論文レビュー] Measuring the Quality of Explanations: The System Causability Scale (SCS). Comparing Human and Machine Explanations
本論文は、説明可能AI(xAI)システムにおける説明の質を評価するための10項目のリッターランク尺度であるSystem Causability Scale(SCS)を紹介する。この尺度は、System Usability Scale(SUS)を模倣して開発され、医療分野におけるAIとの対話インターフェースにおける説明の質を評価することを目的としている。信頼性は高く(Cronbach’s alpha = .91)、Framingham Risk Toolへの応用を通じて検証され、SCSスコアが0.86に達した。これは、説明の因果的透明性(causability)が強く感じられていることを示している。
Recent success in Artificial Intelligence (AI) and Machine Learning (ML) allow problem solving automatically without any human intervention. Autonomous approaches can be very convenient. However, in certain domains, e.g., in the medical domain, it is necessary to enable a domain expert to understand, <i>why</i> an algorithm came up with a certain result. Consequently, the field of Explainable AI (xAI) rapidly gained interest worldwide in various domains, particularly in medicine. Explainable AI studies transparency and traceability of opaque AI/ML and there are already a huge variety of methods. For example with layer-wise relevance propagation relevant parts of inputs to, and representations in, a neural network which caused a result, can be highlighted. This is a first important step to ensure that end users, e.g., medical professionals, assume responsibility for decision making with AI/ML and of interest to professionals and regulators. Interactive ML adds the component of human expertise to AI/ML processes by enabling them to re-enact and retrace AI/ML results, e.g. let them check it for plausibility. This requires new human-AI interfaces for explainable AI. In order to build effective and efficient interactive human-AI interfaces we have to deal with the question of <i>how to evaluate the quality of explanations</i> given by an explainable AI system. In this paper we introduce our System Causability Scale to measure the quality of explanations. It is based on our notion of Causability (Holzinger et al. in Wiley Interdiscip Rev Data Min Knowl Discov 9(4), 2019) combined with concepts adapted from a widely-accepted usability scale.
研究の動機と目的
- 説明可能AI(xAI)システムにおける説明の質を評価するための標準化されたツールの不足に対処すること。
- ドメインエキスパートがAIの説明をどれだけ因果的に理解可能で使いやすいと感じるかを測定する、信頼性が高く、迅速かつ心理学的妥当性を持つインSTRUMENTを開発すること。
- 医療などハイリスク分野における理解、信頼性、意思決定支援の観点から、説明がどれほど効果的に機能するかを評価することで、効果的な人間-AIインターフェースの設計を支援すること。
- 既存の使いやすさ指標と比較してSCSの妥当性を検証し、実世界の医療予測モデルへの適用可能性を示すこと。
提案手法
- 一般的な使いやすさではなく、因果的透明性(causability)に焦点を当てた10項目のリッターランク尺度として、System Usability Scale(SUS)フレームワークを改変した。
- 説明の完全性、文脈内での明確さ、詳細の調整可能性、外部支援への依存性の低さ、説明の即時性という、主な次元を評価する項目を設計した。
- Framingham Risk Tool(FRT)—広く使われている臨床予測モデル—にSCSを適用し、実際の医療現場における説明の質を評価した。
- Cronbach’s alphaを用いて内部的一致性を計算し、元のSUSとの相関関係を測定することで、収束妥当性を評価した。
- 各項目に5段階のリッターランク尺度(1 = 強く反対、5 = 強く賛成)を適用し、合計スコアを0〜1のスケールに正規化した。
- 実際の医療従事者によるパイロット評価を実施し、実世界での使いやすさと説明の質の主観的評価を確認した。
実験結果
リサーチクエスチョン
- RQ1人間-AIインターフェースにおける説明の質を、因果的透明性と使いやすさを反映する形で体系的に測定する方法は何か?
- RQ2SCSは、SUSのような既存の使いやすさ指標とどの程度相関しているか。これは、SCSの妥当性を示唆する。
- RQ3SCSは、Framingham Risk Toolのような実世界の医療AIシステムにおける説明の質の差を効果的に検出できるか?
- RQ4SCSは、異なるユーザーと文脈において、説明の質を測定する上でどの程度信頼性があるか?
主な発見
- System Causability Scale(SCS)は、Cronbach’s alphaが.91に達し、高い内部的一致性を示し、信頼性が非常に高いことが確認された。
- 元のSystem Usability Scale(SUS)との相関係数がr = .985に達し、収束妥当性が強く支持された。
- Framingham Risk Toolに適用した結果、SCSスコアは正規化スケールで0.86に達し、説明の因果的透明性が強く感じられていることが示された。
- 「文脈内で理解可能」という項目(評価5)や「外部支援が不要」という項目(評価5)が最も高いスコアを獲得し、明確さと独立性の高さが裏付けられた。
- 一方で、「知識ベースと併用可能」という項目(評価3)や「素早く習得可能」という項目(評価3)はスコアが低く、インターフェース改善の余地が浮き彫りになった。
- 第二の派生尺度との間でも中程度の相関(r = .664)を示し、SCSが一般的な使いやすさとは別個の構造を測定しているが、関連する概念を捉えていることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。