[논문 리뷰] Exploiting Excessive Invariance caused by Norm-Bounded Adversarial Robustness
이 논문은 ℓp-노름 기반의 노이즈에 강건한 딥 네ural 네트워크를 훈련시키는 과정에서, 의미적 내용은 변화시키지만 예측 결과는 유지하는 invariance-based adversarial examples에 대한 취약성이 간접적으로 증가할 수 있음을 보여준다. 저자들은 이러한 과도한 invariance를 악용하는 model-agnostic 공격을 제안하며, 최신의 robust 모델이 invariance-based adversarial examples에 대해 인간 레이블과 일치하는 비율이 단 38%에 불과함을 입증한다. 이는 기준 모델 대비 유의미하게 낮은 성능이다.
Adversarial examples are malicious inputs crafted to cause a model to misclassify them. Their most common instantiation, "perturbation-based" adversarial examples introduce changes to the input that leave its true label unchanged, yet result in a different model prediction. Conversely, "invariance-based" adversarial examples insert changes to the input that leave the model's prediction unaffected despite the underlying input's label having changed. In this paper, we demonstrate that robustness to perturbation-based adversarial examples is not only insufficient for general robustness, but worse, it can also increase vulnerability of the model to invariance-based adversarial examples. In addition to analytical constructions, we empirically study vision classifiers with state-of-the-art robustness to perturbation-based adversaries constrained by an $\ell_p$ norm. We mount attacks that exploit excessive model invariance in directions relevant to the task, which are able to find adversarial examples within the $\ell_p$ ball. In fact, we find that classifiers trained to be $\ell_p$-norm robust are more vulnerable to invariance-based adversarial examples than their undefended counterparts. Excessive invariance is not limited to models trained to be robust to perturbation-based $\ell_p$-norm adversaries. In fact, we argue that the term adversarial example is used to capture a series of model limitations, some of which may not have been discovered yet. Accordingly, we call for a set of precise definitions that taxonomize and address each of these shortcomings in learning.
연구 동기 및 목표
- 적대적 예외에 대한 강건성이 ℓp-노름 기반의 공격에서 비롯되는지, 의미적으로 의미 있는 입력 변화에 대한 민감도 감소로 이어져 과도한 invariance를 유도하는지 조사하는 것.
- ℓp-강건 모델에서 과도한 invariance가 일반화 능력과 인간 중심의 의사결정을 약화시킨다는 것을 입증하는 것.
- ℓp-balls 내에서 invariance-based adversarial examples를 생성하는 model-agnostic 공격을 제안하고 평가하는 것.
- 적대적 공격에 대한 강건성이 일반적인 강건성을 의미한다는 가정을 도전하며, 적대적 실패 유형의 분류 체계 수립을 주장하는 것.
제안 방법
- ℓp-노름 기반의 공격에 강건하지만 invariance-based 공격에서는 실패하는 분석적 모델을 구축하여, 민감도와 invariance 사이의 상충 관계를 입증하는 것.
- 입력을 의미적으로 의미 있는 방식으로 수정하면서도 ℓp-balls 내에 머무르도록 하여, ℓ0 및 ℓ∞ invariance-based adversarial examples를 생성하는 model-agnostic 공격을 설계하는 것.
- 알고리즘 레이블이 모호한 경우 인간 레이블을 오라클로 사용하여 적대적 예외에서 모델의 일치도를 평가하는 것.
- invariance-based 공격에 대해 state-of-the-art ℓp-robust 비전 분류기 모델을 훈련하고 평가하며, 인간-모델 일치도를 강건성의 대체 지표로 측정하는 것.
- ℓp-노름 제약 조건 하에서 결정 경계의 기하학적 특성을 분석하여, ℓp-노름이 의미적 또는 작업 관련 변화를 포괄하지 못함을 보여주는 것.
- invariance-based 강건성은 perturbation-based 강건성과 별도로 평가되어야 하며, 새로운 지표 및 아키텍처 설계(예: 가역 네트워크)를 통해 invariance를 제어할 것을 제안하는 것.
실험 결과
연구 질문
- RQ1ℓp-노름 기반의 적대적 예외에 대한 강건성이 의미 있는 입력 변화에 대한 민감도 감소로 이어져 과도한 invariance를 유도하는가?
- RQ2ℓp-강건하게 훈련된 모델에 대해 ℓp-balls 내에서 invariance-based adversarial examples를 구성할 수 있으며, 그 효과는 어떠한가?
- RQ3ℓp-강건 모델이 invariance-based adversarial examples에 대해 인간 레이블과 얼마나 일치하는가? 기준 모델 대비 어떤가?
- RQ4과도한 invariance를 유도하는 아키텍처나 훈련 구성 요소가 존재하는가? 그리고 이를 완화할 수 있는가?
- RQ5perturbation-based와 invariance-based 적대적 예외의 조합이 모든 회피 공격 벡터를 포괄하는가, 아니면 다른 실패 유형이 존재하는가?
주요 결과
- ℓp-노름 기반의 적대적 예외에 강건하도록 훈련된 모델은 invariance-based adversarial examples에 대해 인간 레이블과의 일치도가 유의미하게 감소하여, ℓ0 공격의 경우 38%에 불과했으며 이는 기준 모델의 54%에 못 미친다.
- ℓ0 invariance-based 공격은 55%의 경우에서 성공했고, ℓ∞ 공격은 21%의 경우에서 성공했으며, 각각 해당 방어의 ℓp-balls 내에서 수행되었다.
- 성공한 invariance-based adversarial examples에서 ℓp-강건 모델은 ℓ0 공격에 대해 58% 미만, ℓ∞ 공격에 대해 5% 미만의 인간 레이블 일치도를 기록했다.
- 연구 결과에 따르면 일부 클래스 쌍은 다른 클래스 쌍보다 invariance-based 공격에 더 취약한 것으로 나타났으며, 이는 강건성이 데이터셋 전반에 균일하게 분포하지 않음을 시사한다.
- 저자들은 MNIST의 특성에 과도하게 피팅된 방어 기법이 의미 있는 변화에 대해 일반화되지 못할 수 있으며, 이는 높은 ℓ0 및 ℓ∞ 강건성을 주장하는 방어 기법의 문제점을 드러낸다.
- 현재의 평가 프로토콜이 ℓp-balls을 사용하는 데서는 의미 있는 의미 변화와 노름 범위 내의 invariance-based 취약성을 탐지하지 못할 수 있으며, 이는 평가가 부족할 수 있음을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.