Skip to main content
QUICK REVIEW

[論文レビュー] Tell Me a Story! Narrative-Driven XAI with Large Language Models

David Martens, James Hinns|arXiv (Cornell University)|Sep 29, 2023
Explainable Artificial Intelligence (XAI)被引用数 4
ひとこと要約

本論文では、SHAPおよび対向的要因(CF)説明から人間が理解しやすい自然言語の説明を生成するため、大規模言語モデル(LLMs)を活用する物語中心の説明可能AI(XAI)アプローチ、XAIstoriesを紹介する。90%を超える一般ユーザーがSHAPstoriesを説得力があると評価しており、CFstoriesは手作業で書かれた物語に比べて10倍速く、20%以上の精度向上を達成している。これはAIシステムにおける解釈可能性のギャップを埋める強力な解決策を提供する。

ABSTRACT

In many AI applications today, the predominance of black-box machine learning models, due to their typically higher accuracy, amplifies the need for Explainable AI (XAI). Existing XAI approaches, such as the widely used SHAP values or counterfactual (CF) explanations, are arguably often too technical for users to understand and act upon. To enhance comprehension of explanations of AI decisions and the overall user experience, we introduce XAIstories, which leverage Large Language Models to provide narratives about how AI predictions are made: SHAPstories do so based on SHAP explanations, while CFstories do so for CF explanations. We study the impact of our approach on users' experience and understanding of AI predictions. Our results are striking: over 90% of the surveyed general audience finds the narratives generated by SHAPstories convincing. Data scientists primarily see the value of SHAPstories in communicating explanations to a general audience, with 83% of data scientists indicating they are likely to use SHAPstories for this purpose. In an image classification setting, CFstories are considered more or equally convincing as the users' own crafted stories by more than 75% of the participants. CFstories additionally bring a tenfold speed gain in creating a narrative. We also find that SHAPstories help users to more accurately summarize and understand AI decisions, in a credit scoring setting we test, correctly answering comprehension questions significantly more often than they do when only SHAP values are provided. The results thereby suggest that XAIstories may significantly help explaining and understanding AI predictions, ultimately supporting better decision-making in various applications.

研究の動機と目的

  • 複雑なAIモデルのための、人間にとって理解しやすく整合性のある説明のギャップを解消すること、特に解釈可能性が不可欠な重要な分野において。
  • SHAPおよび対向的要因の説明の限界(技術的複雑さや物語的整合性の欠如)を克服し、それらを自然言語の物語に変換すること。
  • 一般ユーザーおよびデータサイエンティストが、LLMによって生成された物語(XAIstories)が手作業で書かれた説明よりも説得力、正確性、効率性において優れているかどうかを評価すること。
  • XAIstoriesが非専門家向けの対象者にAI意思決定を効果的に伝える可能性を評価し、AIシステムにおける信頼性と理解度を向上させること。

提案手法

  • XAIstoriesは、微調整された大規模言語モデル(LLMs)を活用し、予測スコアに与える特徴の寄与度を解釈する物語的説明(SHAPstories)にSHAP値を変換する。
  • 対向的要因の説明に関しては、予測クラスを変えるために必要な入力特徴の最小変更を説明する物語(CFstories)を生成し、「もし~なら」という仮定シナリオとしてフレーミングする。
  • LLMが一貫性があり、事実に基づき、文脈的に適切な物語を生成できるように、SHAPおよびCFの出力から導かれた構造化された入力テンプレートとプロンプト工学を用いる。
  • 本手法は、一般参加者およびデータサイエンティストを対象としたユーザースタディを通じて評価され、LLMによって生成された物語と手作業で書かれた物語、およびベースライン説明との比較が行われる。
  • 評価指標には、説得力の認識、物語の質、作成速度、説明内容の正確性が含まれ、統計的分析により結果の妥当性を検証する。
  • フレームワークは、表形式データ(学生の成績予測)および画像分類(MobileNet V2の誤分類)に適用され、多分野への適用可能性が示された。

実験結果

リサーチクエスチョン

  • RQ1LLMによって生成された物語(XAIstories)は、非専門家ユーザーに対してSHAP説明の明確さと信頼性を顕著に向上させることができるか?
  • RQ2CFstoriesは、手作業で書かれた物語と比較して、物語の質、正確性、作成速度においてどのように異なるか?
  • RQ3データサイエンティストは、XAIstoriesを非技術的ステークホルダーにAI説明を伝えるために、どれほど価値があると認識しているか?
  • RQ4SHAPおよび対向的要因から導かれる物語的説明は、現実世界の意思決定文脈において、ユーザーのAI予測に対する信頼性と理解度を向上させることができるか?

主な発見

  • 一般ユーザーの90%以上がSHAPstoriesを説得力があると評価しており、LLMによって生成された物語がSHAP説明に対して高いユーザー受容性を示している。
  • 92%のデータサイエンティストが、SHAPstoriesによって非専門家がAI予測を理解する際の容易さと自信が向上すると報告した。
  • 83%のデータサイエンティストが、非専門家向けにSHAPstoriesを活用する可能性が非常に高いと回答した。
  • 画像分類の文脈では、75%を超える一般ユーザーがCFstoriesをユーザーが作成した物語と同等またはそれ以上に説得力があると評価した。
  • CFstoriesは、手作業による物語作成時間の10倍の速さで生成され、著しい効率性の向上が示された。
  • CFstoriesは、手作業で作成された説明に比べ、20%以上の正確性向上を達成しており、元の対向的要因の論理に忠実であることが示唆された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。