Skip to main content
QUICK REVIEW

[论文解读] On the Faithfulness Measurements for Model Interpretations.

Fan Yin, Zhouxing Shi|arXiv (Cornell University)|Apr 18, 2021
Topic Modeling参考文献 40被引用 12
一句话总结

本文提出了一套系统化的框架,通过三种标准——基于移除的忠实度、敏感性和稳定性——来评估自然语言处理模型解释的忠实度。该框架引入了受对抗鲁棒性启发的方法,在文本分类和依存句法分析任务中,于所有三项指标上均取得了最先进性能。

ABSTRACT

Recent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions. Despite the surge of new interpretations, it remains an open problem how to define and quantitatively measure the faithfulness of interpretations, i.e., to what extent they conform to the reasoning process behind the model. To tackle these issues, we start with three criteria: the removal-based criterion, the sensitivity of interpretations, and the stability of interpretations, that quantify different notions of faithfulness, and propose novel paradigms to systematically evaluate interpretations in NLP. Our results show that the performance of interpretations under different criteria of faithfulness could vary substantially. Motivated by the desideratum of these faithfulness notions, we introduce a new class of interpretation methods that adopt techniques from the adversarial robustness domain. Empirical results show that our proposed methods achieve top performance under all three criteria. Along with experiments and analysis on both the text classification and the dependency parsing tasks, we come to a more comprehensive understanding of the diverse set of interpretations.

研究动机与目标

  • 为解决后训练阶段自然语言处理模型解释中忠实度的定量测量这一开放性问题。
  • 定义并实现多种不同的忠实度概念,包括基于移除的忠实度、敏感性和稳定性度量。
  • 设计一种系统化的评估范式,以捕捉解释可靠性中的多样化方面。
  • 开发新型解释方法,使其在所有三项忠实度标准上均表现更优。
  • 对文本分类和依存句法分析任务中的解释方法进行全面的实证分析。

提出的方法

  • 提出三种不同的忠实度标准:基于移除的(移除输入标记的影响)、敏感性(对输入微小扰动的响应)和稳定性(在输入变化下的一致性)。
  • 引入一种受对抗鲁棒性技术启发的新解释方法类别,以提升在所有三项标准下的忠实度。
  • 采用一种训练范式,使解释在对抗性扰动下保持鲁棒性,同时保持模型预测的保真度。
  • 将所提出的评估框架应用于文本分类和依存句法分析任务,以确保其广泛适用性。
  • 使用基于梯度和显著性图的解释技术作为基线方法进行对比。
  • 通过消融研究和受控实验,分离并分析每一项忠实度标准的贡献。

实验结果

研究问题

  • RQ1基于移除的忠实度、敏感性和稳定性这三种忠实度标准,在评估解释质量方面有何不同?
  • RQ2现有解释方法在所提出的三项忠实度标准上的表现如何?
  • RQ3对抗鲁棒性技术能否被适配以提升自然语言处理模型解释的忠实度?
  • RQ4所提出的方法是否能同时在所有三项忠实度标准上实现更优性能?
  • RQ5忠实度度量与真实世界自然语言处理任务中的模型性能和可解释性之间存在何种关联?

主要发现

  • 不同解释方法在三项忠实度标准上的表现存在显著差异,表明没有一种方法能在所有忠实度概念上均表现优异。
  • 在某一标准上表现良好的解释(如基于移除的忠实度)往往在其他标准上表现欠佳(如敏感性),凸显了多维度评估的必要性。
  • 所提出的受对抗鲁棒性启发的解释方法在所有三项忠实度标准上均达到最优性能,证明了其有效性和泛化能力。
  • 评估框架成功揭示了现有解释技术中的权衡与局限,提供了对其可靠性更细致的理解。
  • 在文本分类和依存句法分析任务上的实证结果表明,忠实度得到一致提升,验证了所提方法的鲁棒性与可扩展性。
  • 本研究揭示,忠实度并非单一属性,需依赖多种互补度量才能全面评估解释质量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。