[論文レビュー] Explainability Is in the Mind of the Beholder: Establishing the Foundations of Explainable Artificial Intelligence
本稿は、説明可能性を論理的推論を応用した透明なインサイトに適用し、説明対象者の背景知識と文脈的要因を通して解釈することによって、人間中心の説明可能AIの基盤を確立する。説明可能性を単なる知識移転ではなく、理解を促進するインタラクティブなプロセスとして再定義し、透明性対パフォーマンス、およびアンテホック対ポストホック手法に関する長年の議論を解消する統一的フレームワークを提供する。
Explainable artificial intelligence and interpretable machine learning are research domains growing in importance. Yet, the underlying concepts remain somewhat elusive and lack generally agreed definitions. While recent inspiration from social sciences has refocused the work on needs and expectations of human recipients, the field still misses a concrete conceptualisation. We take steps towards addressing this challenge by reviewing the philosophical and social foundations of human explainability, which we then translate into the technological realm. In particular, we scrutinise the notion of algorithmic black boxes and the spectrum of understanding determined by explanatory processes and explainees' background knowledge. This approach allows us to define explainability as (logical) reasoning applied to transparent insights (into, possibly black-box, predictive systems) interpreted under background knowledge and placed within a specific context -- a process that engenders understanding in a selected group of explainees. We then employ this conceptualisation to revisit strategies for evaluating explainability as well as the much disputed trade-off between transparency and predictive power, including its implications for ante-hoc and post-hoc techniques along with fairness and accountability established by explainability. We furthermore discuss components of the machine learning workflow that may be in need of interpretability, building on a range of ideas from human-centred explainability, with a particular focus on explainees, contrastive statements and explanatory processes. Our discussion reconciles and complements current research to help better navigate open questions -- rather than attempting to address any individual issue -- thus laying a solid foundation for a grounded discussion and future progress of explainable artificial intelligence and interpretable machine learning.
研究の動機と目的
- 説明可能AI(XAI)と解釈可能機械学習(IML)におけるコア定義に合意が得られていない問題に対処するため、哲学的および社会科学的原則に基づいてそれらを基盤づける。
- 説明可能性をモデルの静的特性ではなく、理解を促進する動的で文脈依存のプロセスとして再定式化する。
- モデルの透明性と予測性能の間の見かけのトレードオフを解消するため、アンテホック型とポストホック型の説明可能性の違いを明確にし、それぞれのコストと忠実度を分析する。
- 人間の理解と信頼に焦点を当て、XAI技術の設計、評価、導入を支援する概念的フレームワークを提供する。
- 運用文脈とユーザーのニーズに応じて、データ、モデル、予測といった機械学習ワークフローの各要素が、それぞれ解釈可能である必要があることを特定する。
提案手法
- 説明可能性を、予測システムに関する透明なインサイトに論理的推論を適用し、説明対象者の背景知識と運用文脈によって媒介される概念的モデルを提唱する。
- 透明性と解釈可能性が二値的状態ではなく、連続的スケール上に位置する理解のスケールを導入する。
- アンテホック型(本質的に説明可能)とポストホック型(後から適合)の説明可能性技術の違いを分析し、ポストホック手法が本質的に単純または低コストであるとは限らないことを強調する。
- 科学哲学と社会心理学の知見を活用して、説明を一方向の情報伝達ではなく、対話的プロセスとして位置づける。
- ユーザーの期待や質問に合わせて論理的に整合性のある説明を構築する、インタラクティブで物語的アプローチのエージェントとしてのエクスプライヤーを設計することを提言する。
- 単なる事実の提示ではなく、理解を生じさせることのできる説明可能性の評価フレームワークを提唱し、ユーザーのニーズと文脈に適合した混合手法のメトリクスのセットを要請する。
実験結果
リサーチクエスチョン
- RQ1人間の理解を目的とする場合、人工知能における説明可能性とは何か?(知識移転ではなく。)
- RQ2説明対象者の背景知識と運用文脈に依存するプロセスとして、説明可能性をどのように概念化できるか?
- RQ3予測性能と説明可能性の間のトレードオフは、実際の制約であると言えるか? また、アンテホック型とポストホック型の説明手法によってその性質はどのように変化するか?
- RQ4ポストホック型エクスプライヤーは、普遍的で使いやすいと見なされるが、忠実度と信頼性をどのように評価できるか?
- RQ5対照的説明(「なぜこの結果でなく、あの結果か?」という問いに答えるもの)とインタラクティブな対話は、ブラックボックスモデルの理解を深めるために果たす役割は何か?
主な発見
- 説明可能性はモデルそのものの特性ではなく、理解を生じさせる推論、透明なインサイト、文脈的解釈に根ざしたインタラクティブなプロセスである。
- ポストホック型エクスプライヤーの品質は、アンテホック型のものより本質的に低いわけではないが、忠実度と信頼性を確保するためには、顕著な工学的作業と注意深い設計が不可欠である。
- 透明性と予測力の間の見かけのトレードオフは、複雑で文脈依存的であり、どちらか一方を普遍的に優先するルールは存在しない。
- 対照的説明(「なぜこの結果でなく、あの結果か?」)は、特に高リスク分野において、モデル意思決定の挑戦、反論、デバッグを可能にするために不可欠である。
- 公平性、責任性、デバッグを支援するため、説明可能性はデータ、モデル、予測の各段階を含む、機械学習ライフサイクル全体に埋め込まれるべきである。
- XAIの統一的評価フレームワークは、知識移転ではなく理解の促進を最優先とし、ユーザーのニーズと文脈に適合した定性的および定量的メトリクスを統合するべきである。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。