Skip to main content
QUICK REVIEW

[论文解读] Interpretable Deep Learning: Interpretations, Interpretability, Trustworthiness, and Beyond.

Xuhong Li, Haoyi Xiong|arXiv (Cornell University)|Mar 19, 2021
Explainable Artificial Intelligence (XAI)参考文献 137被引用 4
一句话总结

本文通过澄清解释与可解释性等核心概念,提出了一种解释算法的新分类法,并评估其性能与可信度,对可解释深度学习进行了全面综述。此外,还探讨了其与对抗鲁棒性及数据增强的关联,并提供了开源工具以支持实现。

ABSTRACT

Deep neural networks have been well-known for their superb performance in handling various machine learning and artificial intelligence tasks. However, due to their over-parameterized black-box nature, it is often difficult to understand the prediction results of deep models. In recent years, many interpretation tools have been proposed to explain or reveal the ways that deep models make decisions. In this paper, we review this line of research and try to make a comprehensive survey. Specifically, we introduce and clarify two basic concepts-interpretations and interpretability-that people usually get confused. First of all, to address the research efforts in interpretations, we elaborate the design of several recent interpretation algorithms, from different perspectives, through proposing a new taxonomy. Then, to understand the results of interpretation, we also survey the performance metrics for evaluating interpretation algorithms. Further, we summarize the existing work in evaluating models' interpretability using trustworthy interpretation algorithms. Finally, we review and discuss the connections between deep models' interpretations and other factors, such as adversarial robustness and data augmentations, and we introduce several open-source libraries for interpretation algorithms and evaluation approaches.

研究动机与目标

  • 澄清深度学习中‘解释’与‘可解释性’常被混淆的差异。
  • 基于其设计原则与目标,系统性地分类近期的解释算法。
  • 综述并评估用于评估解释算法的性能指标。
  • 研究可信解释方法如何提升模型的可解释性与可靠性。
  • 探究模型解释、对抗鲁棒性与数据增强技术之间的相互作用。

提出的方法

  • 提出一种新分类法,依据其底层机制与目标对解释算法进行分类。
  • 从多个角度回顾并分类各种解释技术,如显著性图、注意力可视化与概念激活。
  • 使用标准化性能指标(包括忠实度、稳定性与保真度)评估解释算法。
  • 分析可信解释在提升模型透明度与用户信心方面的作用。
  • 探讨对抗样本与数据增强对解释质量与模型行为的影响。
  • 推荐并记录开源库,以支持解释方法的实现与评估。

实验结果

研究问题

  • RQ1‘解释’与‘可解释性’的概念有何不同,为何这一区分在深度学习中至关重要?
  • RQ2现代解释算法的关键设计原则与类别是什么,如何系统性地对其进行分类?
  • RQ3哪些性能指标最能有效评估解释方法的质量与可靠性?
  • RQ4可信解释算法如何提升深度模型的可解释性与可信度?
  • RQ5模型解释与对抗鲁棒性、数据增强等因素之间存在何种关系?

主要发现

  • ‘解释’(模型特定的解释)与‘可解释性’(模型固有的可解释性)之间的区别,对严谨评估至关重要。
  • 新提出的解释算法分类法,有助于更清晰地理解并比较不同设计目标下的多样化方法。
  • 忠实度、稳定性与保真度是评估解释输出质量的关键指标。
  • 可信解释方法能显著提升真实应用场景中用户信心与模型可靠性。
  • 对抗鲁棒性与数据增强会影响解释结果的一致性与可靠性。
  • 已有多个开源库可供支持解释技术的实现与评估。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。