Skip to main content
QUICK REVIEW

[论文解读] Why is plausibility surprisingly problematic as an XAI criterion?

Weina Jin, Xiaoxiao Li|arXiv (Cornell University)|Mar 30, 2023
Explainable Artificial Intelligence (XAI)被引用 4
一句话总结

本文主张,在XAI中优化解释的可信度从根本上与模型可理解性、透明度和可信度等核心目标相悖。它表明,可信度正则化会使AI解释过度模仿人类解释,从而削弱真正的解释效用,并可能引发对模型的不信任或过度信任。相反,可信度应仅作为人类理解的中间计算代理,而非最终目标。

ABSTRACT

Explainable artificial intelligence (XAI) is motivated by the problem of making AI predictions understandable, transparent, and responsible, as AI becomes increasingly impactful in society and high-stakes domains. The evaluation and optimization criteria of XAI are gatekeepers for XAI algorithms to achieve their expected goals and should withstand rigorous inspection. To improve the scientific rigor of XAI, we conduct a critical examination of a common XAI criterion: plausibility. Plausibility assesses how convincing the AI explanation is to humans, and is usually quantified by metrics of feature localization or feature correlation. Our examination shows that plausibility is invalid to measure explainability, and human explanations are not the ground truth for XAI, because doing so ignores the necessary assumptions underpinning an explanation. Our examination further reveals the consequences of using plausibility as an XAI criterion, including increasing misleading explanations that manipulate users, deteriorating users' trust in the AI system, undermining human autonomy, being unable to achieve complementary human-AI task performance, and abandoning other possible approaches of enhancing understandability. Due to the invalidity of measurements and the unethical issues, this position paper argues that the community should stop using plausibility as a criterion for the evaluation and optimization of XAI algorithms. We also delineate new research approaches to improve XAI in trustworthiness, understandability, and utility to users, including complementary human-AI task performance.

研究动机与目标

  • 挑战广泛存在的假设,即可信度是评估XAI技术的有效最终目标。
  • 证明优化可信度会损害模型的可理解性、透明度和可信度。
  • 主张可信度应被重新定位为人类理解的中间计算代理,而非主要评估目标。
  • 区分AI解释任务与目标定位任务,强调需要采用面向XAI的专用评估指标。
  • 呼吁XAI研究转向以用户为中心的评估,基于可解释性特定目标和人类推理目标。

提出的方法

  • 分析在XAI中将可信度作为主要评估指标在概念和实践上的缺陷。
  • 将可信度与真实性、忠实度进行对比,指出可信度依赖于人类先验知识和判断,而非模型的正确性。
  • 提出XAI评估应优先考虑对人类推理任务(如决策验证、偏见检测、知识发现)的效用。
  • 区分AI模型任务(如分类)与XAI任务(如解释生成),并主张为两者分别设计评估指标。
  • 建议采用受控实验设计,对内在可解释模型进行测试,以隔离并评估可解释性性能与模型性能。
  • 将可信度定位为人类理解的计算代理,其在优化XAI效用方面具有实用价值,但不应作为独立目标。

实验结果

研究问题

  • RQ1尽管可信度被广泛使用,为何它仍是评估XAI技术的有缺陷的标准?
  • RQ2为何优化可信度会损害模型的透明度和可信度?
  • RQ3在人机交互中,将可信解释等同于正确模型决策会产生何种后果?
  • RQ4如何重构XAI评估体系,以优先考虑可解释性特定效用,而非可信度?
  • RQ5如何区分AI模型性能评估与XAI解释质量评估?

主要发现

  • 对XAI优化可信度会导致解释过度模仿人类解释,无法支持对模型的深层理解或透明度。
  • 基于可信度的优化可能引发对AI模型的不信任或过度信任,因为可信度与模型决策的正确性被解耦。
  • 可信度并非真实性或忠实度的可靠代理;它衡量的是人类的合理性,而非模型的准确性。
  • 可信度可作为中间指标,在XAI效用优化中模拟人类理解,例如在决策验证或偏见检测中具有实用价值。
  • 使用如定位性能等指标(常用于AI模型)来评估XAI系统,会错误地反映XAI性能,并混淆模型任务与解释任务。
  • 本文呼吁实现范式转变,采用面向XAI的评估目标,使其与人类推理目标和可解释性效用对齐,而非仅追求类人可信度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。