[論文レビュー] Don't Explain without Verifying Veracity: An Evaluation of Explainable AI with Video Activity Recognition
本研究では、制御されたユーザ実験を通じて、動画行動認識における説明の真実性が説明可能AI(XAI)に与える影響を評価している。その結果、正確な説明や説明なしの状況と比較して、真実性が低い説明は、ユーザのパフォーマンスと合意度を著しく低下させることを明らかにした。これは、誤解や不信を引き起こすため、誤った説明は全くない状況よりも悪影響を及ぼす可能性があることを示している。
Explainable machine learning and artificial intelligence models have been used to justify a model's decision-making process. This added transparency aims to help improve user performance and understanding of the underlying model. However, in practice, explainable systems face many open questions and challenges. Specifically, designers might reduce the complexity of deep learning models in order to provide interpretability. The explanations generated by these simplified models, however, might not accurately justify and be truthful to the model. This can further add confusion to the users as they might not find the explanations meaningful with respect to the model predictions. Understanding how these explanations affect user behavior is an ongoing challenge. In this paper, we explore how explanation veracity affects user performance and agreement in intelligent systems. Through a controlled user study with an explainable activity recognition system, we compare variations in explanation veracity for a video review and querying task. The results suggest that low veracity explanations significantly decrease user performance and agreement compared to both accurate explanations and a system without explanations. These findings demonstrate the importance of accurate and understandable explanations and caution that poor explanations can sometimes be worse than no explanations with respect to their effect on user performance and reliance on an AI system.
研究の動機と目的
- 説明の真実性が説明可能AIシステムにおけるユーザのパフォーマンスと信頼に与える影響を調査すること。
- 不正確または誤解を招く説明が、動画行動認識タスクにおけるユーザの理解力と意思決定を低下させるかどうかを評価すること。
- 正確な説明、真実性が低い説明、および説明なしの3条件におけるユーザのパフォーマンスと合意度を比較すること。
- 説明の真実性がユーザのメンタルモデルやシステムへの依存度に与える影響を評価すること。
- モデルの正確性を超えて、XAI設計における真実性の重要性を実証的に示すこと。
提案手法
- 120名の参加者を対象に、説明可能動画行動認識システムを用いた制御されたユーザ研究を実施した。
- 正確な説明、真実性が低い(不正確な)説明、および説明なしの3つの実験条件を設計した。
- ユーザのパフォーマンス、システム予測との合意度、正確性の認識を測定するために、動画のレビューとクエリ処理タスクを用いた。
- 構造化されたアンケートとインタラクションログを通じて、ユーザのタスクパフォーマンス、自信、メンタルモデルの構築状況を収集した。
- 多様な非専門家ユーザの代表を確保するため、Amazon Mechanical Turkを用いて参加者を募集した。
- 性能と合意度の違いを比較するために、統計的手法を用いて結果を分析した。
実験結果
リサーチクエスチョン
- RQ1説明の真実性は、動画行動認識タスクにおけるユーザのパフォーマンスにどのように影響するか?
- RQ2不正確な説明を提供すると、正確な説明や説明なしの状況と比較して、ユーザのシステム予測への合意度が低下するか?
- RQ3説明の真実性は、ユーザがシステムの正確性や信頼性をどのように認識するかに影響を与えるか?
- RQ4誤った説明は、ユーザがAIシステムの誤ったメンタルモデルを構築する原因になるか?
- RQ5正確な説明を備えたシステムと説明なしのシステムとの間で、ユーザのパフォーマンスに有意差が生じるか?
主な発見
- 真実性が低い説明は、正確な説明や説明なしの状況と比較して、ユーザのパフォーマンスを著しく低下させた。
- 不正確な説明が提供された際、ユーザのシステム予測への合意度が著しく低下し、信頼性と信頼性の低下が示された。
- 真実性が低い説明に依存する参加者では、モデルの実際の論理と一致する傾向が弱く、誤ったメンタルモデルの構築が示された。
- 本研究では、誤った説明は、完全に説明がない状況よりも有害であることが判明した。誤った説明は混乱や誤解を引き起こした。
- 正確な説明を受けた参加者は、説明なしまたは誤った説明を受けた参加者と比較して、より良い理解力と高いタスクパフォーマンスを示した。
- 結果から、説明の真実性はXAI設計において重要な要因であることが示された。品質の低い説明は、ユーザの効果性と信頼性を損なう可能性がある。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。