Skip to main content
QUICK REVIEW

[論文レビュー] On Formal Feature Attribution and Its Approximation

Jinqiang Yu, Alexey Ignatiev|arXiv (Cornell University)|Jul 7, 2023
Bayesian Modeling and Causal Inference被引用数 5
ひとこと要約

本稿では、特徴量の重要度を、その特徴量が含まれる帰納的説明の割合として定義する、形式的特徴量帰属(FFA)という新規手法を紹介する。これは論理的整合性を保証する。本手法は、いつでも形式的説明を列挙する技術を用いた効率的な近似手法を提案し、表形式および画像データセット、特に実世界のソフトウェア欠険予測において、LIME や SHAP よりも忠実度、ケンダールのtau、RBO の指標で優れた性能を示した。

ABSTRACT

Recent years have witnessed the widespread use of artificial intelligence (AI) algorithms and machine learning (ML) models. Despite their tremendous success, a number of vital problems like ML model brittleness, their fairness, and the lack of interpretability warrant the need for the active developments in explainable artificial intelligence (XAI) and formal ML model verification. The two major lines of work in XAI include feature selection methods, e.g. Anchors, and feature attribution techniques, e.g. LIME and SHAP. Despite their promise, most of the existing feature selection and attribution approaches are susceptible to a range of critical issues, including explanation unsoundness and out-of-distribution sampling. A recent formal approach to XAI (FXAI) although serving as an alternative to the above and free of these issues suffers from a few other limitations. For instance and besides the scalability limitation, the formal approach is unable to tackle the feature attribution problem. Additionally, a formal explanation despite being formally sound is typically quite large, which hampers its applicability in practical settings. Motivated by the above, this paper proposes a way to apply the apparatus of formal XAI to the case of feature attribution based on formal explanation enumeration. Formal feature attribution (FFA) is argued to be advantageous over the existing methods, both formal and non-formal. Given the practical complexity of the problem, the paper then proposes an efficient technique for approximating exact FFA. Finally, it offers experimental evidence of the effectiveness of the proposed approximate FFA in comparison to the existing feature attribution algorithms not only in terms of feature importance and but also in terms of their relative order.

研究の動機と目的

  • LIME や SHAP といったヒューリスティックな特徴量帰属手法が、論理的に整合性がなく、分布外のサンプリングに苦しむという限界を解消すること。
  • 与えられた予測において、特定の特徴量が含まれる帰納的説明の割合として特徴量帰属を形式的に定式化し、論理的整合性を保証すること。
  • 多項式階層の第二レベルに位置する計算複雑性を考慮し、FFA の効率的な近似手法を開発すること。
  • 特徴量帰属の正確さと相対的順位の観点から、既存手法との比較において FFA の性能を評価すること。

提案手法

  • FFA は、与えられた予測における帰納的説明のうち、特定の特徴量を含むものの割合として形式的に定義される。
  • 機械学習モデルの1階論理表現を用いた形式的説明列挙を活用し、特にブースティング木を対象としている。
  • いつでも利用可能な推論アルゴリズムを用い、帰納的説明を段階的に生成することで、FFA のリアルタイム近似を可能にする。
  • 説明の列挙が進むにつれて FFA 値を更新し、収束状態を監視する。
  • 本手法は、表形式および画像データに適用可能であり、特にロジスティック回帰モデルを用いた実世界のソフトウェア欠険予測にも適用した。
  • 評価には標準的な指標(誤差、ケンダールのtau、ランクバイアス付きオーバーラップ(RBO))を用い、LIME や SHAP と比較した。

実験結果

リサーチクエスチョン

  • RQ1形式的特徴量帰属(FFA)は、機械学習モデルにおける特徴量重要度を原理的かつ論理的に整合した測度として定式化可能か?
  • RQ2FFA の近似手法は、忠実度と順位の正確さという観点で、LIME や SHAP といったヒューリスティック手法に比べてどのように性能を発揮するか?
  • RQ3FFA の近似は、速やかに正確な FFA 値に収束するか? これは異なるデータセットやモデルタイプによってどのように変化するか?
  • RQ4LIME や SHAP は、形式的帰属基準とどの程度一致しないのか? その帰属がどのように逸脱しているか?

主な発見

  • 画像データセットにおいて、10秒間の計算後、FFA の近似は LIME や SHAP よりも誤差、ケンダールのtau、RBO のスコアが優れていた。
  • PneumoniaMNIST データセットでは、10秒後、FFA の近似は 0.83 の RBO と 0.66 のケンダールのtauを達成し、LIME や SHAP を上回った。
  • OpenStack および Qt データセットを用いた即時欠険予測において、FFA は LIME や SHAP よりも形式的説明と著しく一致しており、誤差が低く、特徴量順位の一致度が高かった。
  • 本研究では、LIME や SHAP が、多くの形式的説明に含まれる特徴量に対してゼロの重みを割り当てることが多くあり、形式的基準と根本的に不整合であることが判明した。
  • FFA の近似プロセスは急速に収束し、数分以内に高品質な結果が得られた。これは、理論的複雑性を考慮しても実用的妥当性があることを示唆している。
  • 本稿では、LIME や SHAP といったヒューリスティック手法が、形式的帰属基準と一致しないことが実証された。場合によっては、関係のない特徴量に非ゼロの重みを割り当てたり、重要な特徴量を無視したりする場合がある。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。