Skip to main content
QUICK REVIEW

[논문 리뷰] Curls & Whey: Boosting Black-Box Adversarial Attacks

Yucheng Shi, Siyu Wang|arXiv (Cornell University)|2019. 04. 02.
Adversarial Robustness in Machine Learning참고 문헌 31인용 수 13
한 줄 요약

Curls & Whey는 반복적 궤적에서 기울기 상승과 하강을 조합함으로써 전이성과 노이즈 크기를 개선하는 새로운 블랙박스 적대적 공격을 제안한다 (Curls). 또한 강건성 인식 최적화를 통해 노이즈를 정제함으로써 성능을 향상시킨다 (Whey). 이는 ImageNet과 Tiny-ImageNet에서 기존 최고 성능을 넘어서며, ℓ2 노름이 매우 낮게 유지된다. 특히 적대적으로 훈련된 모델과 앙상블 모델에 대해서도 뛰어난 성능을 발휘한다.

ABSTRACT

Image classifiers based on deep neural networks suffer from harassment caused by adversarial examples. Two defects exist in black-box iterative attacks that generate adversarial examples by incrementally adjusting the noise-adding direction for each step. On the one hand, existing iterative attacks add noises monotonically along the direction of gradient ascent, resulting in a lack of diversity and adaptability of the generated iterative trajectories. On the other hand, it is trivial to perform adversarial attack by adding excessive noises, but currently there is no refinement mechanism to squeeze redundant noises. In this work, we propose Curls & Whey black-box attack to fix the above two defects. During Curls iteration, by combining gradient ascent and descent, we `curl' up iterative trajectories to integrate more diversity and transferability into adversarial examples. Curls iteration also alleviates the diminishing marginal effect in existing iterative attacks. The Whey optimization further squeezes the `whey' of noises by exploiting the robustness of adversarial perturbation. Extensive experiments on Imagenet and Tiny-Imagenet demonstrate that our approach achieves impressive decrease on noise magnitude in l2 norm. Curls & Whey attack also shows promising transferability against ensemble models as well as adversarially trained models. In addition, we extend our attack to the targeted misclassification, effectively reducing the difficulty of targeted attacks under black-box condition.

연구 동기 및 목표

  • 기존의 반복적 공격이 기울기 상승을 단지 일관되게 따르는 데서 비롯하는 다양성 부족과 적응성 부족 문제를 해결하고자 한다.
  • 반복 횟수가 증가함에도 불구하고 현재 방법들이 노이즈를 정제하지 못하는 문제를 해결하고자 한다.
  • 적대적 훈련된 모델과 앙상블 모델과 같이 상이한 모델에 대해서도 적대적 예제의 전이성을 향상시키고자 한다.
  • 반복 과정에 보간법을 통합하여 타겟 공격의 효과를 높이고자 한다.

제안 방법

  • Curls 반복은 대체 모델의 손실 함수에 대해 기울기 상승과 하강을 번갈아 적용하여 둥근 궤적을 생성함으로써 다양성을 증가시키고 경계를 더 잘 넘는 데 기여한다.
  • Curls 반복 내에서 이분법 검색을 사용하여 결정 경계에 더 가까운 적대적 예제를 얻고, 이로써 노이즈 크기를 감소시킨다.
  • Whey 최적화는 적대적 편향의 강건성을 활용하여 픽셀을 값별로 그룹화하고, 무작위로 중복된 노이즈를 제거함으로써 노이즈를 정제한다.
  • 타겟 클래스로의 공격을 안내하기 위해 보간법을 반복 과정에 통합함으로써 타겟 분류의 어려움을 줄인다.
  • 질문 제약 조건 하에 대체 모델을 사용하여 적대적 예제를 생성한 후, 반복 후 최적화를 통해 이를 정제한다.

실험 결과

연구 질문

  • RQ1반복 공격에서 기울기 상승과 하강을 조합함으로써 궤적의 다양성과 적대적 예제의 전이성이 향상될 수 있는가?
  • RQ2편향의 강건성을 활용한 정제 메커니즘을 통해 적대적 노이즈를 효과적으로 감소시킬 수 있는가?
  • RQ3반복 공격에 보간법을 통합하면 타겟 공격의 어려움이 뚜렷하게 감소하는가?
  • RQ4Curls & Whey는 적대적 훈련과 모델 앙상블과 같은 강력한 방어 기법에 대해 어떻게 성능을 발휘하는가?

주요 결과

  • Inception-v3를 대상으로 Tiny-ImageNet을 공격할 때 Curls & Whey는 중앙값 ℓ2 노름 2.0633을 기록하며 기준 방법보다 유의미하게 낮게 유지된다.
  • 적대적으로 훈련된 모델에 대해서도 Curls & Whey는 중앙값 ℓ2 노름 2.0633 (Inception-v3)과 2.2852 (Inception-ResNet-v2)을 유지하며 모든 기준 방법을 능가한다.
  • 동일한 질의 예산 하에서 I-FGSM 및 경계 공격 대비 최대 40%까지 노이즈 크기를 감소시킨다.
  • 절단 실험 결과, Curls 반복, 이분법, Whey 최적화 각각의 구성 요소가 노이즈 감소에 기여하는 것으로 확인된다.
  • 보간법 통합으로 타겟 공격이 성공적으로 수행되며, 더 낮은 ℓ2 노름을 기록함으로써 표준 반복 방법 대비 뚜렷한 성능 향상을 보였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.