[论文解读] What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability Methods
本文通过大规模心理物理学实验(n=1,150)在三个真实世界场景中提出了一个人性化评估框架,用于评估可解释性方法:偏见检测、新策略发现和故障案例分析。研究发现,SmoothGrad 是最有效的归因方法,而忠实度度量无法预测人类使用效果,表明可解释性方法可能需要传达‘什么’特征驱动决策,而不仅仅是‘何处’。
A multitude of explainability methods and associated fidelity performance metrics have been proposed to help better understand how modern AI systems make decisions. However, much of the current work has remained theoretical -- without much consideration for the human end-user. In particular, it is not yet known (1) how useful current explainability methods are in practice for more real-world scenarios and (2) how well associated performance metrics accurately predict how much knowledge individual explanations contribute to a human end-user trying to understand the inner-workings of the system. To fill this gap, we conducted psychophysics experiments at scale to evaluate the ability of human participants to leverage representative attribution methods for understanding the behavior of different image classifiers representing three real-world scenarios: identifying bias in an AI system, characterizing the visual strategy it uses for tasks that are too difficult for an untrained non-expert human observer as well as understanding its failure cases. Our results demonstrate that the degree to which individual attribution methods help human participants better understand an AI system varied widely across these scenarios. This suggests a critical need for the field to move past quantitative improvements of current attribution methods towards the development of complementary approaches that provide qualitatively different sources of information to human end-users.
研究动机与目标
- 从人类用户视角评估可解释性方法在真实世界场景中的实际有用性。
- 确定当前的忠实度度量是否准确反映归因方法对终端用户的实用性。
- 调查感知相似性或解释复杂度在多大程度上可预测人类理解模型决策的表现。
- 确定归因方法是否能有效支持人类在复杂或存在偏见的决策情境下理解模型行为。
- 提出一种以人为本的评估框架,超越理论度量,转向实际可用性。
提出的方法
- 在三个真实世界的人工智能决策场景(偏见检测、新策略发现、故障案例理解)中,对 1,150 名参与者开展了大规模心理物理学实验。
- 将人类在预测和决策任务中的表现作为解释有用性的主要衡量标准。
- 计算不同类别之间诊断图像区域的感知相似性得分,以评估感知模糊性是否影响解释的有用性。
- 使用忠实度度量和人工标注的实用性评分,对多种归因方法(如 Grad-CAM、SmoothGrad、Integrated Gradients)进行评估。
- 分析忠实度度量、解释复杂度、感知相似性与实际人类表现之间的相关性,以识别解释失败的预测因素。
- 发布所有数据和代码,以支持可复现性,并促进对以人为本评估框架的未来采用。
实验结果
研究问题
- RQ1在真实世界场景中,现有归因方法在帮助人类用户理解人工智能模型决策方面有多有用?
- RQ2当前的忠实度度量是否能可靠地预测可解释性方法对人类用户的实际有用性?
- RQ3诊断图像区域之间的感知相似性在多大程度上影响了归因图在辅助人类理解方面的有效性?
- RQ4解释复杂度在多大程度上可预测人类在利用解释理解模型决策时的表现?
- RQ5解释中的‘什么’(语义内容)与‘何处’(空间位置)在人类可解释性中分别起到什么作用?
主要发现
- SmoothGrad 在所测试的各类场景中均为最有效的归因方法,显著提升了人类对模型决策的理解能力。
- 忠实度度量与人类估算的实用性之间无显著相关性,表明其作为实际有用性预测指标表现极差。
- 诊断图像区域之间的感知相似性得分,比忠实度或复杂度度量更能有效预测解释失败。
- 解释复杂度与人类表现之间仅存在微弱相关性,表明其并非解释实用性中的主要影响因素。
- 当不同类别之间的诊断特征在感知上相似时(如不同品种的猫与狗),无论忠实度或复杂度如何,归因方法均无法有效支持人类理解。
- 结果表明,当前的归因方法在支持人类理解方面存在根本性局限,因其未传达‘什么’特征驱动了决策,而仅说明了‘何处’。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。