Skip to main content
QUICK REVIEW

[論文レビュー] Interpreting Adversarial Examples by Activation Promotion and Suppression

Kaidi Xu, Sijia Liu|arXiv (Cornell University)|Apr 3, 2019
Adversarial Robustness in Machine Learning参考文献 33被引用数 39
ひとこと要約

論文はピクセル、画像、ネットワークの視点から敵対的摂動を分析し、活性化における促進-抑制効果(PSE)を明らかにし、摂動をクラス識別領域と意味概念に結びつける。さらに解釈性主導のロバスト性とニューロンマスキングを防御として探究。

ABSTRACT

It is widely known that convolutional neural networks (CNNs) are vulnerable to adversarial examples: images with imperceptible perturbations crafted to fool classifiers. However, interpretability of these perturbations is less explored in the literature. This work aims to better understand the roles of adversarial perturbations and provide visual explanations from pixel, image and network perspectives. We show that adversaries have a promotion-suppression effect (PSE) on neurons' activations and can be primarily categorized into three types: i) suppression-dominated perturbations that mainly reduce the classification score of the true label, ii) promotion-dominated perturbations that focus on boosting the confidence of the target label, and iii) balanced perturbations that play a dual role in suppression and promotion. We also provide image-level interpretability of adversarial examples. This links PSE of pixel-level perturbations to class-specific discriminative image regions localized by class activation mapping (Zhou et al. 2016). Further, we examine the adversarial effect through network dissection (Bau et al. 2017), which offers concept-level interpretability of hidden units. We show that there exists a tight connection between the units' sensitivity to adversarial attacks and their interpretability on semantic concepts. Lastly, we provide some new insights from our interpretation to improve the adversarial robustness of networks.

研究の動機と目的

  • ピクセルレベルの摂動がピクセルおよび画像の観点からCNN予測にどのように影響するかを研究する。
  • 注意機構と領域ベースの分析を用いて異なる敵対的攻撃の有効性を説明する。
  • 敵対的摂動が内部CNN表現と概念検出器に与える影響を調べる。
  • 摂動パターンをクラス識別画像領域や意味概念に結びつけ、ロバスト性の洞察を得る。

提案手法

  • 摂動領域の真のクラスとターゲットクラスのロジットの変化を集約するピクセルレベル感度測度を定義する。
  • Promotion-Suppression Ratio (PSR)を導入し、摂動を抑制支配・促進支配・バランス支配に分類する。
  • クラス活性化マップ(CAM)を用いて摂動が識別的な画像領域にどのように整合するかを解釈する。
  • 識別的領域への摂動の整合性を定量化するためにCAMに基づく解釈可能性スコア(IS)を適用する。
  • ネットワーク分解を用いて摂動が概念検出器とその解釈可能性にどのように影響するかを評価する。
  • CAMを用いて摂動を制約し有効性を研究する改良戦略を提案する。

実験結果

リサーチクエスチョン

  • RQ1ピクセルレベルの摂動は真のラベルとターゲットラベルのいずれに対してCNNの活性化を促進または抑制するか?
  • RQ2CAMとISは敵対的摂動の画像レベルの解釈性をどのように定量化できるか?
  • RQ3ネットワーク分解によって明らかになる内部の概念と敵対的摂動の関係はどうなるか?
  • RQ4解釈ガイド付きの改良とニューロンマスキングはクリーン精度を大幅に損なうことなくロバスト性を向上させることができるか?
  • RQ5摂動パターンは異なる攻撃とターゲットに対して識別的な画像領域とどのように関連するか?

主な発見

  • 敵対的摂動はニューロンの活性化に促進-抑制効果を示し、抑制支配・促進支配・バランス支配のいずれかとして分類される。
  • CAMベースの分析は摂動が真のラベルおよびターゲットラベルの識別領域と整合し、攻撃の画像レベルの解釈性を可能にする。
  • ピクセル摂動と意味概念との間にはネットワーク分解を通じて測定可能な関連があり、感度の高いユニットはしばしば解釈可能な概念である。
  • 解釈性主導の改良は高感度領域を狙うことでより効果的な攻撃を生み出す可能性がある一方、特定の CAM 制約付き改良はより大きな摂動を必要とする場合がある。
  • 敵対的摂動に対して感度の高いユニットは、特に深いネットワーク層で高い解釈性を持つ概念検出器に対応する傾向がある。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。