Skip to main content
QUICK REVIEW

[論文レビュー] Wasserstein Smoothing: Certified Robustness against Wasserstein Adversarial Attacks

Alexander Levine, Soheil Feizi|arXiv (Cornell University)|Oct 23, 2019
Adversarial Robustness in Machine Learning参考文献 22被引用数 16
ひとこと要約

本稿では、ピクセル輸送の流れ空間におけるランダム化スムージングを用いて、ワッサーシュタイン敵対的攻撃に対する最初の認証可能防御であるワッサーシュタインスムージングを提案する。画像の差を流れとして表現し、この空間におけるL1ノルムを用いてワッサーシュタイン距離の上界を定めることで、認証可能ロバスト性を達成し、MNISTおよびCIFAR-10において、保護されていないモデルと比較して顕著に向上した実験的ロバスト性を示した。

ABSTRACT

In the last couple of years, several adversarial attack methods based on different threat models have been proposed for the image classification problem. Most existing defenses consider additive threat models in which sample perturbations have bounded L_p norms. These defenses, however, can be vulnerable against adversarial attacks under non-additive threat models. An example of an attack method based on a non-additive threat model is the Wasserstein adversarial attack proposed by Wong et al. (2019), where the distance between an image and its adversarial example is determined by the Wasserstein metric ("earth-mover distance") between their normalized pixel intensities. Until now, there has been no certifiable defense against this type of attack. In this work, we propose the first defense with certified robustness against Wasserstein Adversarial attacks using randomized smoothing. We develop this certificate by considering the space of possible flows between images, and representing this space such that Wasserstein distance between images is upper-bounded by L_1 distance in this flow-space. We can then apply existing randomized smoothing certificates for the L_1 metric. In MNIST and CIFAR-10 datasets, we find that our proposed defense is also practically effective, demonstrating significantly improved accuracy under Wasserstein adversarial attack compared to unprotected models.

研究の動機と目的

  • 非加法的でワッサーシュタイン距離に基づく敵対的攻撃に対する認証可能防御の欠如に応えること。
  • 正規化されたピクセル強度間の地球移動距離で測定される摂動を想定するワッサーシュタイン脅威モデル下での分類器に対するロバスト性証明を確立すること。
  • ワッサーシュタイン距離を流れ表現空間におけるL1距離に変換することで、既存のL1スムージング証明を可能にすること。
  • 画像分類ベンチマークにおいて、ワッサーシュタイン攻撃下でのベースラインモデルと比較して、提案手法の実験的ロバスト性が優れていることを実証的に検証すること。

提案手法

  • 画像差の非一意な流れ表現を導入し、最小流れのL1ノルムが画像間のワッサーシュタイン距離に等しくなるようにする。
  • この流れ空間表現を活用して、既存のL1ベースのランダム化スムージング証明を適用し、認証可能ロバスト性を実現する。
  • ノイズを流れ空間に追加して分類器をスムージングし、スムージング済み分類器のロバスト性を流れ摂動のL1ノルムから導出する。
  • 多チャンネル画像において、各色チャンネルに独立してノイズを追加することで、付録の補題2によるロバスト性証明を保持する。
  • 1入力あたり10,000回のノイズサンプルを用いたモンテカルロサンプリングにより、スムージング済み分類器の予測を推定し、認証可能ロバスト性を計算する。
  • 実験的ロバスト性は、128回のノイズサンプルを用いた勾配の推定で、ワッサーシュタイン距離に基づく投影勾配降下法を用いて評価する。

実験結果

リサーチクエスチョン

  • RQ1エアス・ムーバー距離に基づく非加法的攻撃であるワッサーシュタイン敵対的攻撃に対して、認証可能防御を構築することは可能か?
  • RQ2流れ表現空間においてワッサーシュタイン距離をL1距離に変換することで、既存のL1スムージング証明を有効化することは可能か?
  • RQ3提案手法の実験的ロバスト性は、保護されていないモデルおよび他の防御手法と比較して、ワッサーシュタイン攻撃下でどのように異なるか?
  • RQ4ワッサーシュタイン脅威モデル下で、認証可能ロバスト性と実験的ロバスト性の差はどの程度か?

主な発見

  • MNISTでは、σ=0.00005で分類精度87.01%、放棄率0.24%を達成し、ワッサーシュタイン攻撃下で保護されていないベースラインと顕著に優れていた。
  • CIFAR-10では、σ=0.0002で精度77.57%、放棄率0.66%を達成し、ワッサーシュタイン攻撃に対して強い実験的ロバスト性を示した。
  • 中央値としての認証可能ロバスト性半径は、MNISTで0.000223、CIFAR-10で0.000179であり、認証可能半径は中央値の実験的攻撃半径と比べて2桁小さいことがわかった。
  • Wongら(2019)が検証した標準的およびバイナリ化モデルと比較して、CIFAR-10において実験的ロバスト性で優れていたが、大きな摂動に対しては adversarially trained モデルほどではなかった。
  • 認証可能ロバスト性は実験的ロバスト性に比べて顕著に低く、ワッサーシュタイン距離に基づくより強力な攻撃や、改善された証明の必要性が示された。
  • 色画像に対しても、各チャンネルごとに流れ空間に独立してノイズを追加することで、本手法は有効であり、付録の補題2が理論的妥当性を裏付けている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。