Skip to main content
QUICK REVIEW

[论文解读] Connecting Algorithmic Research and Usage Contexts: A Perspective of Contextualized Evaluation for Explainable AI

Q. Vera Liao, Yunfeng Zhang|arXiv (Cornell University)|Jun 22, 2022
Explainable Artificial Intelligence (XAI)被引用 8
一句话总结

本文通过识别在不同使用情境(如模型调试、决策支持和审计)中评估标准相对重要性的变化,提出了情境化的可解释人工智能(XAI)评估方法。通过对XAI专家和众包工作者的调查,发现诸如不确定性传达和透明度等标准在实践中备受重视,但在当前的算法研究中却未得到足够重视,呼吁转向以用户为中心的评估,基于真实世界的目标。

ABSTRACT

Recent years have seen a surge of interest in the field of explainable AI (XAI), with a plethora of algorithms proposed in the literature. However, a lack of consensus on how to evaluate XAI hinders the advancement of the field. We highlight that XAI is not a monolithic set of technologies -- researchers and practitioners have begun to leverage XAI algorithms to build XAI systems that serve different usage contexts, such as model debugging and decision-support. Algorithmic research of XAI, however, often does not account for these diverse downstream usage contexts, resulting in limited effectiveness or even unintended consequences for actual users, as well as difficulties for practitioners to make technical choices. We argue that one way to close the gap is to develop evaluation methods that account for different user requirements in these usage contexts. Towards this goal, we introduce a perspective of contextualized XAI evaluation by considering the relative importance of XAI evaluation criteria for prototypical usage contexts of XAI. To explore the context dependency of XAI evaluation criteria, we conduct two survey studies, one with XAI topical experts and another with crowd workers. Our results urge for responsible AI research with usage-informed evaluation practices, and provide a nuanced understanding of user requirements for XAI in different usage contexts.

研究动机与目标

  • 为解决算法化XAI研究与真实世界使用情境之间的脱节问题,即评估标准往往与用户需求不一致。
  • 识别并优先考虑对特定XAI使用情境(如模型调试、决策支持和审计)最为重要的评估标准。
  • 通过实证研究调查终端用户与XAI专家在XAI标准认知上的差异,揭示当前评估实践中的盲点。
  • 通过提供情境特定的权重,指导从业者和研究人员选择或优化XAI技术。
  • 倡导负责任的AI研究,强调基于应用场景、以用户为导向的评估,而非纯粹的算法指标。

提出的方法

  • 通过综合现有文献,构建了XAI评估标准的分类体系,包括忠实性、稳定性、可理解性、不确定性传达和透明度等。
  • 基于用户目标,识别出XAI的典型使用情境,如模型调试、决策支持、能力评估和审计。
  • 开展两项调查研究:一项针对XAI领域专家(HCI与AI领域),另一项针对作为AI应用终端用户的众包工作者。
  • 使用回顾性意见调查,对每种使用情境中评估标准的相对重要性进行排序,生成情境特定的重要性权重。
  • 提出一种加权评估框架,使用情境特定的权重对XAI技术进行评分,以指导选择与优化。
  • 将结果应用于基准测试现有XAI技术,并识别在特定使用场景中应优先考虑的评估标准。

实验结果

研究问题

  • RQ1在模型调试、决策支持和审计等不同使用情境中,XAI评估标准的相对重要性如何变化?
  • RQ2终端用户与XAI专家在评估标准(如忠实性、可理解性和不确定性传达)的重要性认知上有多大程度的一致性?
  • RQ3哪些评估标准对支持特定用户目标(如学习某一领域或评估模型可靠性)最为关键?
  • RQ4如何将评估指标从纯粹的算法属性重新定向为反映真实世界需求的以用户为中心的标准?
  • RQ5在当前XAI研究与评估实践中忽视不确定性传达和透明度等标准,会产生什么影响?

主要发现

  • 不确定性传达和透明度在专家和终端用户中均被列为高度重要的标准,但在当前的XAI评估框架中却显著被忽视。
  • 在决策支持情境中,可理解性和不确定性传达被优先考虑,甚至超过忠实性,表明用户存在权衡取舍。
  • 在模型调试情境中,忠实性和稳定性被评定为更重要,这与算法研究的关注点一致,但可能因此忽略了其他情境中的用户需求。
  • 专家认知与终端用户优先级之间存在显著差距,尤其是在个性化和交互性方面,表明当前评估设计存在盲点。
  • 研究发现忠实性并非始终被优先考虑;在能力评估等情境中,用户愿意为更好的可理解性而牺牲忠实性。
  • 研究结果为加权评估框架提供了基础,使从业者能够根据情境特定的标准重要性选择或优化XAI技术。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。