Skip to main content
QUICK REVIEW

[论文解读] (Un)reasonable Allure of Ante-hoc Interpretability for High-stakes Domains: Transparency Is Necessary but Insufficient for Comprehensibility

Kacper Sokol, Julia E. Vogt|arXiv (Cornell University)|Jun 4, 2023
Explainable Artificial Intelligence (XAI)被引用 4
一句话总结

本文批判了在医疗等高风险领域中对事后可解释性的过度依赖,认为透明度本身不足以实现真正可理解的解释。文章提出了一种以人为中心的框架,整合模块化、溯源性、演化路径和推理类型,以确保各类受众都能真正理解解释内容。

ABSTRACT

Ante-hoc interpretability has become the holy grail of explainable artificial intelligence for high-stakes domains such as healthcare; however, this notion is elusive, lacks a widely-accepted definition and depends on the operational context. It can refer to predictive models whose structure adheres to domain-specific constraints, or ones that are inherently transparent. The latter conceptualisation assumes observers who judge this quality, whereas the former presupposes them to have technical and domain expertise (thus alienating other groups of explainees). Additionally, the distinction between ante-hoc interpretability and the less desirable post-hoc explainability, which refers to methods that construct a separate explanatory model, is vague given that transparent predictive models may still require (post-)processing to yield suitable explanatory insights. Ante-hoc interpretability is thus an overloaded concept that comprises a range of implicit properties, which we unpack in this paper to better understand what is needed for its safe adoption across high-stakes domains. To this end, we outline modelling and explaining desiderata that allow us to navigate its distinct realisations in view of the envisaged application and audience.

研究动机与目标

  • 解决高风险领域(如医疗)中事后可解释性概念模糊且缺乏共识的问题。
  • 阐明为何内在透明的模型仍可能无法被非专家用户理解。
  • 提出一组结构化的设计要求(模块化、溯源性、演化路径和推理类型),以超越透明度提升可理解性。
  • 通过聚焦解释性洞察的来源与处理方式,明确区分事后可解释性与事后解释性。
  • 指导可解释人工智能(XAI)系统的设计,使其不仅透明,而且对各类受众具有实际意义的可理解性。

提出的方法

  • 通过功能、结构和受众导向的标准,将事后可解释性与事后方法区分开来,厘清其概念内涵。
  • 提出基于四个关键设计要求的框架:解释的模块化、解释信息的溯源性(内生性与外生性)、洞察的演化路径(透明度)以及推理类型(人类、算法或混合型)。
  • 分析解释性洞察是源于模型结构(内生性)还是代理模型(外生性),强调其真实性与可追溯性。
  • 评估用于解释模型输出的推理过程,如规则解释、反事实生成和系数分析,以衡量其认知负荷与解释准确性。
  • 提出一种基于谱系的可解释性方法,将技术沿溯源性–演化路径轴进行映射,区分高透明度的事后方法与低透明度的事后方法。
  • 通过医疗案例和决策树的实例,说明模型透明度并不保证用户理解,尤其当推理被错误归因或过度简化时。

实验结果

研究问题

  • RQ1为何事后可解释性常被认为更优,尽管其定义缺乏稳定且广泛接受的共识?
  • RQ2在高风险领域中,模型透明度在多大程度上能确保非专家用户理解?
  • RQ3解释性洞察的溯源性与演化路径在多大程度上影响解释的可信度与可靠性?
  • RQ4人类、算法或混合型推理在将模型输出转化为有意义解释的过程中起到何种作用?
  • RQ5如何设计可解释人工智能系统,以确保解释不仅透明,而且对多样化受众具有可理解性?

主要发现

  • 事后可解释性是一个被过度使用且模糊的概念,常将透明度与可理解性混淆,导致高风险领域中产生错位的期望。
  • 即使模型本身具有内在透明性,非专家用户仍可能因认知错配而无法理解,例如错误地将决策树的根节点分裂视为特征重要性的指标。
  • 直接从透明模型中提取的内生性解释,其溯源可靠性与透明度高于依赖代理模型的外生性解释,后者存在准确性风险。
  • 解释背后的推理过程(人类、算法或混合型)显著影响其可理解性,必须与解释对象的专业水平相匹配。
  • 解释的溯源性与演化路径对信任至关重要,内生性洞察相比事后代理模型具有更高的透明度与更低的失真风险。
  • 必须采用以人为中心的方法,整合模块化设计、面向特定受众的推理方式以及可追溯的信息演化路径,才能弥合透明度与实际可理解性之间的鸿沟。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。