Skip to main content
QUICK REVIEW

[论文解读] A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems

Sina Mohseni, Niloofar Zarei|arXiv (Cornell University)|Nov 28, 2018
Explainable Artificial Intelligence (XAI)被引用 12
一句话总结

本文提出了一种多学科框架,用于设计和评估可解释人工智能(XAI)系统,通过在机器学习、可视化和人机交互领域对设计目标和评估方法进行分类。该框架提供了一套分步、迭代的指南与评估技术,旨在统一跨学科的XAI研究,并支持端到端的系统开发。

ABSTRACT

The need for interpretable and accountable intelligent systems grows along with the prevalence of artificial intelligence applications used in everyday life. Explainable intelligent systems are designed to self-explain the reasoning behind system decisions and predictions, and researchers from different disciplines work together to define, design, and evaluate interpretable systems. However, scholars from different disciplines focus on different objectives and fairly independent topics of interpretable machine learning research, which poses challenges for identifying appropriate design and evaluation methodology and consolidating knowledge across efforts. To this end, this paper presents a survey and framework intended to share knowledge and experiences of XAI design and evaluation methods across multiple disciplines. Aiming to support diverse design goals and evaluation methods in XAI research, after a thorough review of XAI related papers in the fields of machine learning, visualization, and human-computer interaction, we present a categorization of interpretable machine learning design goals and evaluation methods to show a mapping between design goals for different XAI user groups and their evaluation methods. From our findings, we develop a framework with step-by-step design guidelines paired with evaluation methods to close the iterative design and evaluation cycles in multidisciplinary XAI teams. Further, we provide summarized ready-to-use tables of evaluation methods and recommendations for different goals in XAI research.

研究动机与目标

  • 通过统一设计与评估方法,解决XAI研究在不同学科间分散的问题。
  • 识别并分类面向不同XAI用户群体(如终端用户、领域专家、开发者)的特定设计目标。
  • 将评估方法与具体设计目标相映射,以支持系统化且有针对性的XAI系统开发。
  • 开发一种实用的、迭代的框架,以在多学科XAI团队中闭合设计与评估循环。
  • 提供即用型表格,列出针对特定XAI研究目标的评估方法与建议。

提出的方法

  • 对机器学习、可视化和人机交互(HCI)领域的XAI文献进行全面调查。
  • 根据目标用户群体(包括终端用户、领域专家和开发者)对XAI设计目标进行分类。
  • 根据评估方法的关注重点,将其划分为计算型、以用户为中心和系统级评估三类。
  • 提出一种嵌套的、迭代的框架,将多层中的设计目标与适当的评估方法相连接。
  • 整合交互式机器学习和可用性评估领域现有框架的见解,以指导系统设计与评估层级。
  • 通过为期一年的多学科案例研究对框架进行验证,参与的八位研究人员具备多样化专业背景。

实验结果

研究问题

  • RQ1如何在不同用户群体和学科之间系统性地对XAI设计目标进行分类?
  • RQ2针对特定XAI设计目标,哪些评估方法最为合适,且其适用性如何随用户群体而变化?
  • RQ3如何通过统一框架有效支持XAI系统设计中的跨学科协作?
  • RQ4在XAI系统中,如何有效对齐计算可解释性与人类可解释性?
  • RQ5如何构建迭代设计与评估循环,以确保XAI系统具备鲁棒性并符合用户需求?

主要发现

  • 建立了XAI设计目标与评估方法之间的清晰映射,揭示了机器学习、可视化和HCI领域在优先事项上的显著差异。
  • 该框架通过将内层的算法可解释性与外层的用户体验及系统结果相连接,成功支持了迭代设计与评估。
  • 案例研究证明了该框架在指导多学科团队完成为期一年的XAI开发过程中的实际效用。
  • 评估方法因目标用户而异:计算指标用于模型可解释性评估,而基于受试者的人类实验则对评估用户信任与理解至关重要。
  • 该框架具备可扩展性和可扩展性,可与特定领域设计指南集成,并可通过更多案例研究进一步验证。
  • 局限性在于抽象层次较高,可能需要补充针对特定领域的详细实现指导。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。