[논문 리뷰] Adversarial purification with Score-based generative models
이 논문은 denoising score-matching 훈련된 energy-based model을 사용한 adversarial purification을 도입하여, 공격당한 이미지를 빠르고 강인하게 정화하고 certified robustness를 위한 무작위 노이즈 강화까지 가능하게 한다.
While adversarial training is considered as a standard defense method against adversarial attacks for image classifiers, adversarial purification, which purifies attacked images into clean images with a standalone purification model, has shown promises as an alternative defense method. Recently, an Energy-Based Model (EBM) trained with Markov-Chain Monte-Carlo (MCMC) has been highlighted as a purification model, where an attacked image is purified by running a long Markov-chain using the gradients of the EBM. Yet, the practicality of the adversarial purification using an EBM remains questionable because the number of MCMC steps required for such purification is too large. In this paper, we propose a novel adversarial purification method based on an EBM trained with Denoising Score-Matching (DSM). We show that an EBM trained with DSM can quickly purify attacked images within a few steps. We further introduce a simple yet effective randomized purification scheme that injects random noises into images before purification. This process screens the adversarial perturbations imposed on images by the random noises and brings the images to the regime where the EBM can denoise well. We show that our purification method is robust against various attacks and demonstrate its state-of-the-art performances.
연구 동기 및 목표
- 적대적 정화를 adversarial training과는 별개의 실질적인 방어책으로 제시한다.
- denoising score matching을 활용하여 정화를 위한 score-based 모델을 훈련한다.
- 적응형 단계 크기를 갖는 결정적 업데이트 정화 방식 개발.
- 무작위 노이즈 주입으로 강건성 강화 및 randomized smoothing을 통해 인증된 강건성 입증.
- 표준 데이터셋에서 강력한 adaptive attacks에 대해 최첨단 성능을 시연한다.
제안 방법
- 정화를 위한 score function을 학습하기 위해 denoising score matching (dsm)을 사용하여 energy-based model을 훈련한다.
- 다중 스케일 Noise Conditional Score Network (NCSN)을 사용하여 서로 다른 섭동 수준을 처리한다.
- Langevin dynamics가 아닌 방식으로 학습된 score에 의해 안내되는 결정적 업데이트로 정화를 수행한다.
- 정화 전에 무작위 Gaussian 노이즈 주입을 도입하여 강건성을 향상시키고 randomized smoothing을 가능하게 한다.
- 정화 중에 단계 크기를 적응시켜 score 표면에서의 통제된 강하를 보장한다.
- 여러 번의 randomized purification 실행을 결합하고 출력들을 앙상블하여 최종 예측에 사용한다.
실험 결과
연구 질문
- RQ1denoising score matching으로 훈련된 score-based 모델이 전통적인 MCMC 기반 EBM보다 더 빠르게 적대적 예제를 효과적으로 정화할 수 있는가?
- RQ2정화 전에 무작위 노이즈 주입이 norm-bounded 위협 모델 하에서 인증된 강건성을 산출하는가?
- RQ3제안된 방법이 강력한 adaptive attacks에 대해 기존의 adversarial purification 방법들과 비교하여 어떤 성능을 보이나?
- RQ4정확도와 정화 속도 측면에서 결정적 정화 업데이트가 확률적 Langevin dynamics보다 바람직한가?
- RQ5적응형 단계 크기가 광범위한 하이퍼파라미터 튜닝 없이도 정화의 안정성을 향상시킬 수 있는가?
주요 결과
- dsm-based EBMs를 이용한 결정적 score-가이드 정화는 이전의 장시간 MCMC 접근법보다 공격된 이미지를 훨씬 빠르게 정화한다.
- 정화 전에 무작위 Gaussian 노이즈를 주입하면 강건성이 개선되고 randomized smoothing이 가능해져 특정 위협 모델 하에서 인증된 강건성을 달성한다.
- 이 방법은 여러 adaptive 공격에 대해 강한 강건성을 달성하며 CIFAR-10/100 및 기타 데이터셋에서 기존의 adversarial purification 및 adversarial training 기반선과 비교해 우수한 성능을 보인다.
- DSM으로 학습된 다중 스케일 score 네트워크 (NCSN)가 다양한 공격 강도와 데이터 섭동에 걸쳐 정화를 향상시킨다.
- 적응형 단계 크기 전략이 정화를 추가로 안정시키고 공격 하에서 최종 분류 정확도를 향상시킨다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.