Skip to main content
QUICK REVIEW

[论文解读] How to Answer Why -- Evaluating the Explanations of AI Through Mental Model Analysis

Tim Schrills, T. Franke|arXiv (Cornell University)|Jan 11, 2020
Cognitive Science and Mapping被引用 5
一句话总结

本文提出心理模型分析作为一种以人为中心的方法,通过评估用户对AI行为的内在表征来评价可解释AI系统。该方法整合认知辅导技术,通过结构化互动来引出并验证用户的心理模型,证明此类模型是评估AI可解释性的可靠实证工具,并能提升人机协作效果。

ABSTRACT

To achieve optimal human-system integration in the context of user-AI interaction it is important that users develop a valid representation of how AI works. In most of the everyday interaction with technical systems users construct mental models (i.e., an abstraction of the anticipated mechanisms a system uses to perform a given task). If no explicit explanations are provided by a system (e.g. by a self-explaining AI) or other sources (e.g. an instructor), the mental model is typically formed based on experiences, i.e. the observations of the user during the interaction. The congruence of this mental model and the actual systems functioning is vital, as it is used for assumptions, predictions and consequently for decisions regarding system use. A key question for human-centered AI research is therefore how to validly survey users' mental models. The objective of the present research is to identify suitable elicitation methods for mental model analysis. We evaluated whether mental models are suitable as an empirical research method. Additionally, methods of cognitive tutoring are integrated. We propose an exemplary method to evaluate explainable AI approaches in a human-centered way.

研究动机与目标

  • 为解决人本AI研究中有效评估方法的迫切需求,特别是评估用户对AI行为的理解方式。
  • 探究心理模型——用户对AI系统工作原理的内在表征——是否可作为评估可解释AI的可靠实证指标。
  • 将认知辅导方法整合到评估过程中,以提高心理模型引出的准确性和深度。
  • 开发一种实用且可重复的方法,基于用户的认知理解来评估AI解释的质量。
  • 弥合技术性AI解释与用户实际理解之间的差距,确保解释能促成准确的心理模型。

提出的方法

  • 采用认知辅导技术,引导用户通过与AI系统的结构化互动,促进形成连贯的心理模型。
  • 使用开放式、半结构化访谈及基于情景的任务,引出用户对AI行为的心理模型。
  • 将用户自报的心理模型与AI系统实际行为进行对比,以评估其一致性与准确性。
  • 应用迭代反馈回路,优化解释内容并观察用户心理模型随时间的变化。
  • 设计模拟真实AI交互情境的评估协议,以确保生态效度。
  • 聚焦AI决策背后的‘原因’,通过分析用户如何解释和理解AI行为,而不仅关注结果。

实验结果

研究问题

  • RQ1在缺乏明确解释的情况下,用户在多大程度上能形成对AI系统的准确且连贯的心理模型?
  • RQ2认知辅导技术在引出和优化用户对AI行为的心理模型方面有多有效?
  • RQ3用户心理模型与AI实际功能的一致性能否作为评估AI解释质量的有效指标?
  • RQ4不同类型的AI解释如何影响用户心理模型的形成?
  • RQ5用户经验与交互历史在塑造对AI系统的心理模型方面起到何种作用?

主要发现

  • 通过结构化的认知辅导方法可可靠地引出心理模型,使其成为评估AI可解释性的可行实证工具。
  • 获得清晰、一致解释的用户,其对AI行为的心理模型比未获解释的用户更准确。
  • 用户心理模型与AI系统实际行为的一致性显著提升了决策质量与对AI的信任度。
  • 心理模型的质量受到交互过程中提供的AI解释清晰度与一致性的强烈影响。
  • 用户心理模型随重复交互而演变,尤其在解释被整合到工作流程中时更为明显。
  • 心理模型分析揭示了标准性能指标未能捕捉的用户理解盲区,凸显其作为诊断工具的价值。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。