[论文解读] Mind the box: $l_1$-APGD for sparse adversarial attacks on image classifiers
该论文提出 $l_1$-APGD,一种新型对抗攻击方法,通过在 $l_1$-球与图像域 $[0,1]^d$ 的交集中进行正确投影,显著提升了图像分类器的稀疏对抗攻击性能,而此前的 $l_1$-PGD 方法忽略了这一约束。该方法在对抗训练中实现了最先进的 $l_1$-鲁棒性,并构建了 $l_1$-AutoAttack,一种可靠的集成攻击方法,在计算开销极小的情况下,优于先前方法,能更有效地评估 $l_1$-鲁棒性。
We show that when taking into account also the image domain $[0,1]^d$, established $l_1$-projected gradient descent (PGD) attacks are suboptimal as they do not consider that the effective threat model is the intersection of the $l_1$-ball and $[0,1]^d$. We study the expected sparsity of the steepest descent step for this effective threat model and show that the exact projection onto this set is computationally feasible and yields better performance. Moreover, we propose an adaptive form of PGD which is highly effective even with a small budget of iterations. Our resulting $l_1$-APGD is a strong white-box attack showing that prior works overestimated their $l_1$-robustness. Using $l_1$-APGD for adversarial training we get a robust classifier with SOTA $l_1$-robustness. Finally, we combine $l_1$-APGD and an adaptation of the Square Attack to $l_1$ into $l_1$-AutoAttack, an ensemble of attacks which reliably assesses adversarial robustness for the threat model of $l_1$-ball intersected with $[0,1]^d$.
研究动机与目标
- 解决现有 $l_1$-PGD 攻击方法在威胁模型中忽略图像域约束 $[0,1]^d$ 所导致的次优性问题。
- 设计一种计算高效的精确投影方法,实现对 $l_1$-球与 $[0,1]^d$ 交集的投影,实验证明该方法可产生更强的攻击效果。
- 设计一种自适应、无需调参的 PGD 方案($l_1$-APGD),在无需人工调整学习率的情况下提升攻击性能。
- 构建 $l_1$-AutoAttack,一种由 $l_1$-APGD、目标 $l_1$-FAB 和 $l_1$-Square Attack 组成的集成攻击,用于可靠且高效的 $l_1$-鲁棒性评估。
提出的方法
- 提出一种精确投影算子 $P_S(u)$,用于投影到集合 $S = B_1(x, \nabla L(x^{(i)})) \cap [0,1]^d$,该方法计算可行且比近似投影更精确。
- 推导出在约束威胁模型 $S$ 下的正确最速下降方向,从而实现更有效的扰动更新。
- 提出 $l_1$-APGD,一种完全自适应的 PGD 变体,可动态选择步长而无需用户输入,从而提升鲁棒性与收敛性。
- 将 Square Attack 适配至 $l_1$-范数,并将其整合入 $l_1$-AutoAttack,通过融合多种攻击类型实现全面的鲁棒性评估。
- 采用多阶段攻击策略,结合交叉熵损失与目标 DLR 损失,并通过多次随机重启提升攻击成功率。
实验结果
研究问题
- RQ1为何标准 $l_1$-PGD 攻击在图像分类任务中表现次优,尽管其专为 $l_1$-扰动设计?
- RQ2能否高效计算 $l_1$-球与 $[0,1]^d$ 交集的精确投影?该投影是否能产生更强的对抗攻击?
- RQ3与固定步长的 PGD 相比,自适应且无需调参的 PGD 方案是否能提升攻击性能与鲁棒性评估效果?
- RQ4$l_1$-APGD 是否能与其它攻击方法有效结合,形成可靠的集成攻击(如 $l_1$-AutoAttack),以评估 $l_1$-鲁棒性?
- RQ5在对抗训练中使用 $l_1$-APGD 是否能训练出具有最先进 $l_1$-鲁棒准确率的模型?
主要发现
- 对 $S = B_1(x,\epsilon) \cap [0,1]^d$ 的精确投影在计算上是可行的,且相比近似投影,显著提升了攻击性能。
- $l_1$-APGD 在 CIFAR-100($\epsilon=12$)上实现了最高的 $l_1$-鲁棒准确率,优于标准 PGD 及其他基线方法。
- 在 ImageNet($\epsilon=60$)上,$l_1$-APGD 在 $l_2$-鲁棒模型上达到 40.5% 的鲁棒准确率,在 $l_\infty$-鲁棒模型上达到 4.4%,与或超过其他攻击方法的表现。
- $l_1$-AutoAttack 在 ImageNet 和 CIFAR-100 的 4 个实验中的 2 个案例中实现了最低的鲁棒准确率,表明其能可靠地估计模型鲁棒性。
- $l_1$-APGD 攻击速度显著快于 B&B 方法,仅需 254 秒处理 1000 张图像(100 步),而 B&B 需要 3612 秒。
- $l_1$-AutoAttack 集成攻击始终优于单一攻击方法,为 $l_1$-鲁棒性评估提供了可靠的基准。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。