Skip to main content
QUICK REVIEW

[論文レビュー] Mind the box: $l_1$-APGD for sparse adversarial attacks on image classifiers

Francesco Croce, Matthias Hein|arXiv (Cornell University)|Mar 1, 2021
Adversarial Robustness in Machine Learning被引用数 7
ひとこと要約

本稿では、$l_1$-ballと画像ドメイン $[0,1]^d$ の共通部分に正確に射影する手法を導入することで、従来の $l_1$-PGD手法が無視していた制約を考慮した、新たな adversarial attack 手法 $l_1$-APGD を提案する。この手法は、 adversarial training における $l_1$-robustness を最先端水準にまで向上させ、計算コストの増加を最小限に抑えつつ、信頼性の高いアンサンブル攻撃 $l_1$-AutoAttack を可能にする。

ABSTRACT

We show that when taking into account also the image domain $[0,1]^d$, established $l_1$-projected gradient descent (PGD) attacks are suboptimal as they do not consider that the effective threat model is the intersection of the $l_1$-ball and $[0,1]^d$. We study the expected sparsity of the steepest descent step for this effective threat model and show that the exact projection onto this set is computationally feasible and yields better performance. Moreover, we propose an adaptive form of PGD which is highly effective even with a small budget of iterations. Our resulting $l_1$-APGD is a strong white-box attack showing that prior works overestimated their $l_1$-robustness. Using $l_1$-APGD for adversarial training we get a robust classifier with SOTA $l_1$-robustness. Finally, we combine $l_1$-APGD and an adaptation of the Square Attack to $l_1$ into $l_1$-AutoAttack, an ensemble of attacks which reliably assesses adversarial robustness for the threat model of $l_1$-ball intersected with $[0,1]^d$.

研究の動機と目的

  • 攻撃の脅威モデルにおいて画像ドメイン制約 $[0,1]^d$ を無視する既存の $l_1$-PGD 攻撃の非最適性を是正すること。
  • 計算的に実行可能な正確な射影を、$l_1$-ball と $[0,1]^d$ の共通部分に定義し、これがより強い攻撃をもたらすことを示すこと。
  • ユーザーによるステップサイズの手動チューニングを必要としない、完全に適応的な PGD 方式($l_1$-APGD)を設計し、攻撃性能と収束性を向上させること。
  • $l_1$-APGD、$l_1$-FAB、$l_1$-Square Attack の3つの攻撃を組み合わせたアンサンブルである $l_1$-AutoAttack を構築し、$l_1$-robustness の信頼性の高い評価を可能とすること。

提案手法

  • 集合 $S = B_1(x, \nabla L(x^{(i)})) \cap [0,1]^d$ への正確な射影演算子 $P_S(u)$ を提案する。これは計算的に実行可能であり、近似射影よりも精度が高く、より優れた攻撃性能を実現する。
  • 制約付き脅威モデル $S$ における正しい勾配降下方向を導出することで、より効果的な摂動更新が可能になる。
  • $l_1$-APGD を導入する。これはユーザー入力なしでステップサイズを動的に選択する完全に適応的な PGD 変種であり、robustness と収束性の両方を向上させる。
  • Square Attack を $l_1$-ノルムに適応させ、$l_1$-AutoAttack に統合することで、複数の攻撃タイプを組み合わせ、包括的な robustness 評価を実現する。
  • 交差エントロピー損失とターゲット付き DLR 損失を用いたマルチフェーズ攻撃戦略を採用し、複数回のランダムリスタートを活用して攻撃成功率を向上させる。

実験結果

リサーチクエスチョン

  • RQ1標準的な $l_1$-PGD 攻撃は、$l_1$-摂動を想定しているにもかかわらず、なぜ画像分類において非最適なのであろうか?
  • RQ2$l_1$-ball と $[0,1]^d$ の共通部分への正確な射影は、計算的に実行可能であり、より強い adversarial 攻撃をもたらすのであろうか?
  • RQ3固定ステップサイズの PGD と比較して、適応的でパrameter-free な PGD 方式は、攻撃性能と robustness 評価を向上させるのであろうか?
  • RQ4$l_1$-APGD を他の攻撃と効果的に組み合わせることで、$l_1$-AutoAttack のような信頼性の高いアンサンブルを構築できるのであろうか?
  • RQ5$l_1$-APGD を adversarial training に用いることで、最先端の $l_1$-robust accuracy を達成するモデルが得られるのであろうか?

主な発見

  • 集合 $S = B_1(x,\epsilon) \cap [0,1]^d$ への正確な射影は、計算的に実行可能であり、近似射影と比較して攻撃性能が顕著に向上することが示された。
  • CIFAR-100($\epsilon=12$)において、$l_1$-APGD は $\epsilon=12$ の条件下で最高の $l_1$-robust accuracy を達成し、標準的な PGD や他のベースラインを上回った。
  • ImageNet($\epsilon=60$)において、$l_1$-APGD は $l_2$-robust モデルで 40.5%、$l_\infty$-robust モデルで 4.4% の robust accuracy を達成し、他の攻撃と同等またはそれを上回った。
  • $l_1$-AutoAttack は、ImageNet と CIFAR-100 の4つのケースのうち2つにおいて、最小の robust accuracy を記録しており、モデルの robustness を信頼性高く推定できることを示している。
  • $l_1$-APGD 攻撃は B&B よりも著しく高速であり、1000枚の画像に対して100ステップで254秒で実行可能であるのに対し、B&B は3612秒を要した。
  • $l_1$-AutoAttack アンサンブルは、個々の攻撃を常に上回り、$l_1$-robustness 評価の信頼性の高いベンチマークを提供している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。