Skip to main content
QUICK REVIEW

[論文レビュー] On the Faithfulness Measurements for Model Interpretations.

Fan Yin, Zhouxing Shi|arXiv (Cornell University)|Apr 18, 2021
Topic Modeling参考文献 40被引用数 12
ひとこと要約

本稿では、除去に基づく忠実性、感受性、安定性の3つの基準を用いて、NLPモデルの解釈の忠実性を体系的に評価するフレームワークを提案する。敵対的ロバスト性にインspiredされた手法を導入し、テキスト分類および依存解析タスクにおいて、3つの指標すべてで最先端の性能を達成する。

ABSTRACT

Recent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions. Despite the surge of new interpretations, it remains an open problem how to define and quantitatively measure the faithfulness of interpretations, i.e., to what extent they conform to the reasoning process behind the model. To tackle these issues, we start with three criteria: the removal-based criterion, the sensitivity of interpretations, and the stability of interpretations, that quantify different notions of faithfulness, and propose novel paradigms to systematically evaluate interpretations in NLP. Our results show that the performance of interpretations under different criteria of faithfulness could vary substantially. Motivated by the desideratum of these faithfulness notions, we introduce a new class of interpretation methods that adopt techniques from the adversarial robustness domain. Empirical results show that our proposed methods achieve top performance under all three criteria. Along with experiments and analysis on both the text classification and the dependency parsing tasks, we come to a more comprehensive understanding of the diverse set of interpretations.

研究の動機と目的

  • 後処理によるNLPモデル解釈の忠実性を定量的に測定するという未解決問題に対処すること。
  • 除去に基づく、感受性、安定性の各指標を含む、忠実性の複数の明確な概念を定義・実装すること。
  • 解釈の信頼性の多様な側面を捉える体系的な評価パラダイムを構築すること。
  • 3つの忠実性基準すべてで性能を向上させる新しい解釈手法を開発すること。
  • テキスト分類および依存解析タスクにおける解釈手法の包括的な実証的分析を提供すること。

提案手法

  • 除去に基づく忠実性(入力トークンの除去が与える影響)、感受性(微小な入力摂動への反応)、安定性(入力の変化に対する一貫性)の3つの異なる忠実性基準を提案する。
  • 敵対的ロバスト性技術にインspiredされた、3基準すべての忠実性を向上させる新しい解釈手法のクラスを導入する。
  • 敵対的摂動に対して解釈がロバストであるように最適化すると同時に、モデルの予測の忠実性を保持する訓練パラダイムを採用する。
  • 広範な適用性を確保するため、提案された評価フレームワークをテキスト分類および依存解析タスクに適用する。
  • 勾配ベースおよびサリエンシーに基づく解釈手法を比較のためのベースラインとして使用する。
  • アブレーションスタディおよび制御実験を用いて、各忠実性基準の貢献を分離する。

実験結果

リサーチクエスチョン

  • RQ1除去に基づく忠実性、感受性、安定性の各基準は、解釈品質の評価においてどのように異なるか?
  • RQ2既存の解釈手法は、提案された3つの忠実性基準のどの程度を満たしているか?
  • RQ3敵対的ロバスト性技術は、NLPモデル解釈の忠実性を向上させるために適応可能か?
  • RQ4提案された手法は、3つの忠実性基準すべてで優れた性能を同時に達成できるか?
  • RQ5忠実性指標は、現実世界のNLPタスクにおけるモデル性能および解釈可能性とどの程度相関しているか?

主な発見

  • 解釈手法の性能は3つの忠実性基準において顕著に異なることが判明し、どの手法も忠実性のあらゆる側面で優れているわけではないことが示された。
  • ある基準(例:除去に基づく忠実性)で優れた性能を発揮する解釈は、他の基準(例:感受性)では劣ることが多く、多面的な評価の必要性が浮き彫りになった。
  • 提案された敵対的ロバスト性にインspiredされた解釈手法は、3基準すべてで最高の性能を達成し、その有効性と一般化能力を示した。
  • 評価フレームワークは、既存の解釈手法におけるトレードオフや限界を効果的に明らかにし、信頼性に関するより洗練された理解を可能にした。
  • テキスト分類および依存解析における実証的結果から、忠実性の向上が一貫して得られ、提案手法のロバスト性とスケーラビリティが裏付けられた。
  • 忠実性は単一の性質ではなく、解釈品質を包括的に評価するには、複数の補完的指標が必要であることが明らかになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。