Skip to main content
QUICK REVIEW

[论文解读] Explainability Is in the Mind of the Beholder: Establishing the Foundations of Explainable Artificial Intelligence

Kacper Sokol, Peter Flach|arXiv (Cornell University)|Dec 29, 2021
Explainable Artificial Intelligence (XAI)被引用 13
一句话总结

本文通过将可解释性定义为应用于透明洞察的逻辑推理,并通过解释对象的背景知识和情境因素进行解读,为以人为本的可解释人工智能(XAI)奠定了基础。它将可解释性重新构认为一种促进理解而非单纯知识传递的互动过程,提出了一个统一的框架,用于评估可解释性并解决长期存在的关于透明性与性能、事前(ante-hoc)与事后(post-hoc)方法之间的争议。

ABSTRACT

Explainable artificial intelligence and interpretable machine learning are research domains growing in importance. Yet, the underlying concepts remain somewhat elusive and lack generally agreed definitions. While recent inspiration from social sciences has refocused the work on needs and expectations of human recipients, the field still misses a concrete conceptualisation. We take steps towards addressing this challenge by reviewing the philosophical and social foundations of human explainability, which we then translate into the technological realm. In particular, we scrutinise the notion of algorithmic black boxes and the spectrum of understanding determined by explanatory processes and explainees' background knowledge. This approach allows us to define explainability as (logical) reasoning applied to transparent insights (into, possibly black-box, predictive systems) interpreted under background knowledge and placed within a specific context -- a process that engenders understanding in a selected group of explainees. We then employ this conceptualisation to revisit strategies for evaluating explainability as well as the much disputed trade-off between transparency and predictive power, including its implications for ante-hoc and post-hoc techniques along with fairness and accountability established by explainability. We furthermore discuss components of the machine learning workflow that may be in need of interpretability, building on a range of ideas from human-centred explainability, with a particular focus on explainees, contrastive statements and explanatory processes. Our discussion reconciles and complements current research to help better navigate open questions -- rather than attempting to address any individual issue -- thus laying a solid foundation for a grounded discussion and future progress of explainable artificial intelligence and interpretable machine learning.

研究动机与目标

  • 通过将可解释人工智能(XAI)和可解释机器学习(IML)的核心定义建立在哲学和社会科学原理之上,解决该领域内核心定义缺乏共识的问题。
  • 将可解释性重新构认为不是模型的静态属性,而是一种动态的、依赖情境的过程,旨在促进人类解释对象的理解。
  • 通过区分事前(ante-hoc)和事后(post-hoc)可解释性,并分析其各自的成本与保真度,解决模型透明性与预测性能之间看似存在的权衡。
  • 提供一个概念性框架,指导XAI技术的设计、评估与部署,重点关注人类的理解与信任。
  • 识别机器学习工作流中的各个组成部分——数据、模型、预测——根据操作情境和用户需求,可能各自需要可解释性。

提出的方法

  • 提出可解释性作为对预测系统透明洞察应用逻辑推理的概念模型,由解释对象的背景知识和操作情境所中介。
  • 引入理解的谱系,将透明性与可解释性视为连续统中的不同位置,而非二元状态。
  • 分析事前(内在可解释)与事后(事后增强)可解释性技术之间的区别,强调事后方法并非本质上更简单或成本更低。
  • 借鉴科学哲学与社会心理学的洞见,将解释视为对话过程,而非单向信息传递。
  • 建议将解释器设计为交互式、叙事驱动的代理,构建与用户期望和问题相匹配的逻辑一致的解释。
  • 提出一种基于可解释性能否激发理解而非仅传递事实陈述的评估框架,并呼吁采用一套混合方法度量指标,以评估其有效性。

实验结果

研究问题

  • RQ1当目标是人类理解而非知识传递时,人工智能中的可解释性由什么构成?
  • RQ2可解释性如何被构想为一种依赖于解释对象背景知识和操作情境的过程?
  • RQ3可解释性与预测性能之间的权衡在多大程度上是真实存在的约束?这种权衡在事前与事后解释技术之间如何变化?
  • RQ4鉴于事后解释器被认为具有普适性和易用性,如何评估其保真度与可靠性?
  • RQ5对比性解释与交互式对话在促进对黑箱模型的深入理解方面发挥什么作用?

主要发现

  • 可解释性并非仅属于模型的属性,而是一种根植于推理、透明洞察与情境化解读的互动过程,最终促成人类解释对象的理解。
  • 事后解释器的质量并不天然低于事前解释器,但其开发需要大量工程投入与精心设计,以确保保真度与可信度。
  • 透明性与预测能力之间的权衡是微妙且依赖情境的,不存在普遍偏好其中一方的规则。
  • 对比性解释——回答“为何是此结果而非彼结果?”——对于使用户能够挑战、反驳或调试模型决策至关重要,尤其在高风险领域。
  • 可解释性必须贯穿整个机器学习生命周期,包括数据、模型和预测阶段,以支持公平性、问责制与调试。
  • XAI的统一评估框架必须优先考虑理解而非知识传递,并整合针对用户需求与情境量身定制的定性与定量度量指标。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。