Skip to main content
QUICK REVIEW

[논문 리뷰] Towards Defending Multiple $\ell_p$-norm Bounded Adversarial Perturbations via Gated Batch Normalization

Aishan Liu, Shiyu Tang|arXiv (Cornell University)|2020. 12. 03.
Adversarial Robustness in Machine Learning참고 문헌 50인용 수 5
한 줄 요약

이 논문은 깊이 있는 신경망의 다중 $\ell_p$-노름으로 제한된 적대적 공격에 대한 강건성을 향상시키는 새로운 정규화 기법인 게이트드 배치 정규화(Gated Batch Normalization, GBN)를 제안한다. 배치 정규화 레이어에 학습 가능한 게이트(합성곱 및 완전 연결)를 통합함으로써 GBN은 특징 표현을 적응적으로 조절하며, CIFAR-10과 Tiny-ImageNet에서 $\ell_1$, $\ell_2$, $\ell_\infty$ 공격에 대해 최신 기술 수준의 강건한 정확도를 달성한다. 정상 정확도는 83.6%이며, AutoAttack 기준 강건 정확도는 최대 75.6%에 이른다.

ABSTRACT

There has been extensive evidence demonstrating that deep neural networks are vulnerable to adversarial examples, which motivates the development of defenses against adversarial attacks. Existing adversarial defenses typically improve model robustness against individual specific perturbation types (\eg, $\ell_{\infty}$-norm bounded adversarial examples). However, adversaries are likely to generate multiple types of perturbations in practice (\eg, $\ell_1$, $\ell_2$, and $\ell_{\infty}$ perturbations). Some recent methods improve model robustness against adversarial attacks in multiple $\ell_p$ balls, but their performance against each perturbation type is still far from satisfactory. In this paper, we observe that different $\ell_p$ bounded adversarial perturbations induce different statistical properties that can be separated and characterized by the statistics of Batch Normalization (BN). We thus propose Gated Batch Normalization (GBN) to adversarially train a perturbation-invariant predictor for defending multiple $\ell_p$ bounded adversarial perturbations. GBN consists of a multi-branch BN layer and a gated sub-network. Each BN branch in GBN is in charge of one perturbation type to ensure that the normalized output is aligned towards learning perturbation-invariant representation. Meanwhile, the gated sub-network is designed to separate inputs added with different perturbation types. We perform an extensive evaluation of our approach on commonly-used dataset including MNIST, CIFAR-10, and Tiny-ImageNet, and demonstrate that GBN outperforms previous defense proposals against multiple perturbation types (\ie, $\ell_1$, $\ell_2$, and $\ell_{\infty}$ perturbations) by large margins.

연구 동기 및 목표

  • 기존 방어 기법이 특정 $\ell_p$-노름 공격에 대해서만 효과를 보이며 동시에 여러 유형에 대해선 성능이 떨어지는 한계를 해결하기 위해.
  • PGD, C&W, AutoAttack 및 자연스러운 변형 공격을 포함한 다양한 적대적 공격에 대한 모델의 강건성을 향상시키기 위해.
  • 다른 $\ell_p$-노름 위협 모델에 일반화할 수 있는 통합된 방어 메커니즘을 개발하기 위해.
  • GBN 블록 내에서 게이트 유형(합성곱 대 완전 연결)과 배치 방식이 강건성과 일반화에 미치는 영향을 조사하기 위해.

제안 방법

  • 학습 가능한 게이트(합성곱 또는 완전 연결)를 배치 정규화 레이어에 통합하여, 각 레이어별로 배치 통계를 적응적으로 조절하는 게이트드 배치 정규화(Gated Batch Normalization, GBN)를 제안함.
  • 불변성과 표현력의 균형을 고려해 하이브리드 게이트 전략을 적용: 첫 번째 GBN 블록에는 합성곱 게이트를, 이후 블록에는 완전 연결 게이트를 사용함.
  • $\ell_1$, $\ell_2$, $\ell_\infty$ 노름 기반의 PGD 공격을 사용한 적대적 훈련을 수행하며, BN은 평가 모드로 유지함.
  • 적대적 데이터의 다양성을 확보하기 위해 훈련 중에 모든 네 개의 BN 브랜치를 통해 기울기 흐름을 유도함으로써 일반화 성능 향상.
  • 청결한 데이터와 다중 노름 기반의 적대적 예제를 조합하여 모델을 훈련하며, 각 공격 유형에 맞게 하이퍼파라미터를 조정함.
  • MNIST, CIFAR-10, Tiny-ImageNet에서 PGD, C&W, BBA, MI-FGSM, SPSA, NATTACK, AutoAttack 등 포괄적인 공격 세트를 대상으로 GBN의 성능을 평가함.

실험 결과

연구 질문

  • RQ1통합된 정규화 기법이 동시에 여러 $\ell_p$-노름으로 제한된 적대적 공격에 대응할 수 있는가?
  • RQ2GBN 블록 내에서 게이트 유형(합성곱 대 완전 연결)과 배치 방식이 강건성과 일반화에 어떤 영향을 미치는가?
  • RQ3적대적 훈련 중에 모든 BN 브랜치를 통해 기울기 흐름을 유도하는 것이 방어 성능 향상에 기여하는가?
  • RQ4AutoAttack 및 C&W와 같은 강력한 공격에 대해 GBN은 기존 방어 기법들(예: AVG, MAX, MSD, PAT, TRADES)과 비교해 어떻게 성능을 내는가?
  • RQ5GBN는 다양한 공격 유형에 걸쳐 높은 청결 정확도를 유지하면서도 강력한 강건성을 확보할 수 있는가?

주요 결과

  • WideResNet-28-10에서 GBN은 83.6%의 청결 정확도와 AutoAttack 기준 75.6%의 강건 정확도를 달성하여 모든 기준 모델을 초월함.
  • CIFAR-10에서 GBN은 BBA 공격 기준 71.1%의 강건 정확도, PGD-$\ell_2$ 공격 기준 70.4%의 강건 정확도를 기록하며, PAT, TRADES 및 기타 방어 기법을 뛰어넘음.
  • GBN 블록 내에서 하이브리드 게이트 전략(합성곱 + 완전 연결)은 PGD-$\ell_1$ 공격 기준 59.6%의 강건 정확도를 기록하며, '모든 합성곱'(30.4%) 및 '모든 완전 연결'(35.1%) 전략보다 뚜렷이 뛰어남.
  • GBN는 가우스 노이즈 및 SPSA와 같은 자연스러운 공격에 대해서도 강력한 성능을 유지하며, WideResNet-28-10 기준 각각 77.4% 및 70.7%의 강건 정확도 기록.
  • ResNet-20에서 GBN은 C&W-$\ell_2$ 공격 기준 74.9%의 강건 정확도를 달성하여, 동일 공격 조건에서 PAT(50.9%) 및 TRADES(59.5%)를 뛰어넘음.
  • WideResNet-28-10에서 PGD-$\ell_\infty$ 공격 기준 GBN은 60.2%의 정확도를 기록하며, MAX(44.1%) 및 AVG(49.6%) 전략보다 뛰어남.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.