Skip to main content
QUICK REVIEW

[論文レビュー] Exploring the Trade-off between Plausibility, Change Intensity and Adversarial Power in Counterfactual Explanations using Multi-objective Optimization

Javier Del Ser, Alejandro Barredo-Arrieta|arXiv (Cornell University)|May 20, 2022
Adversarial Robustness in Machine Learning被引用数 4
ひとこと要約

本稿では、GANを用いたデータ分布モデルと多目的最適化ソルバーを用いて、妥当性、変化強度、敵対的パワーのバランスをとった、複数の目的を最適化する枠組みを提案する。この枠組みにより、画像および3Dデータにおいて、より信頼性が高く、直感的で、実行可能である対応反転解釈が得られ、非専門家ユーザー向けのバイアス検出とモデルの解釈可能性の向上が可能になる。

ABSTRACT

There is a broad consensus on the importance of deep learning models in tasks involving complex data. Often, an adequate understanding of these models is required when focusing on the transparency of decisions in human-critical applications. Besides other explainability techniques, trustworthiness can be achieved by using counterfactuals, like the way a human becomes familiar with an unknown process: by understanding the hypothetical circumstances under which the output changes. In this work we argue that automated counterfactual generation should regard several aspects of the produced adversarial instances, not only their adversarial capability. To this end, we present a novel framework for the generation of counterfactual examples which formulates its goal as a multi-objective optimization problem balancing three different objectives: 1) plausibility, i.e., the likeliness of the counterfactual of being possible as per the distribution of the input data; 2) intensity of the changes to the original input; and 3) adversarial power, namely, the variability of the model's output induced by the counterfactual. The framework departs from a target model to be audited and uses a Generative Adversarial Network to model the distribution of input data, together with a multi-objective solver for the discovery of counterfactuals balancing among these objectives. The utility of the framework is showcased over six classification tasks comprising image and three-dimensional data. The experiments verify that the framework unveils counterfactuals that comply with intuition, increasing the trustworthiness of the user, and leading to further insights, such as the detection of bias and data misrepresentation.

研究の動機と目的

  • 非専門家ユーザーのモデルの説明可能性の欠落を補うために、直感的で実行可能な対応反転解釈を生成すること。
  • 対応反転解釈生成を、妥当性、変化強度、敵対的パワーのバランスをとる多目的最適化問題として形式化すること。
  • 実世界のデータ分布を反映した対応反転解釈を生成することで、モデルの信頼性を高めること。
  • データの誤表現や敵対的摂動に対する脆弱性を含め、モデル挙動のより深い洞察を提供すること。

提案手法

  • 入力データ分布をモデル化するために、生成対抗ネットワーク(GAN)を訓練し、妥当な対応反転解釈の生成を可能にする。
  • GANの識別器を属性ベクトルで条件づけることで、望ましい出力変化を持つ対応反転解釈の生成をガイドする。
  • 妥当性、変化強度、敵対的パワーの3つの目的をバランスとる対応反転解釈を求めるために、多目的最適化ソルバーを採用する。
  • 妥当性は、元のデータ分布に近い実際のサンプルを生成できるGANの能力によって評価される。
  • 変化強度は、元の入力と対応反転解釈との間のL2距離として測定される。
  • 敵対的パワーは、対応反転解釈をターゲットモデルに入力した際の出力変化の大きさによって定量化される。

実験結果

リサーチクエスチョン

  • RQ1対応反転解釈生成は、対立する目標を含む、本質的に多目的最適化問題であるか?
  • RQ2生成された対応反転解釈は、データおよびタスクの文脈に基づく直感的な期待に適合しているか?
  • RQ3多基準対応反転解釈は、トレーニングセット内の隠れたバイアスやデータの誤表現を明らかにできるか?
  • RQ4妥当性、変化強度、敵対的パワーのトレードオフは、ユーザーの信頼性と解釈可能性にどのように影響するか?

主な発見

  • 6つの実験におけるパレート最適解の近似から、対応反転解釈生成が、しばしば対立する複数の目的に支配されており、多目的最適化が不可欠であることが確認された。
  • 生成された対応反転解釈は、色の変化や構造的強調といった妥当な変化を示し、人間の直感とタスク固有の意味論と整合的であった。
  • このフレームワークは、特に属性-クラス関係において、構成的バイアスやデータの誤表現を効果的に特定し、潜在的な脆弱性を示した。
  • 敵対的パワーが高く、変化強度が低い対応反転解釈は、より実行可能であることが判明した一方で、すべてのデータセットで高い妥当性が維持された。
  • 非専門家ユーザーが、モデルの挙動をよりよく理解できるようになるため、直感的で実行可能な代替解釈を提供できるようになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。