Skip to main content
QUICK REVIEW

[논문 리뷰] The curse of overparametrization in adversarial training: Precise analysis of robust generalization for random features regression

Hamed Hassani, Adel Javanmard|arXiv (Cornell University)|2022. 01. 13.
Adversarial Robustness in Machine Learning인용 수 4
한 줄 요약

이 논문은 고차원 점근적 조건 하에서 적대적 훈련을 받는 랜덤 특징 모델의 강건한 일반화에 대한 정밀한 이론적 분석을 제공한다. 과다 매개변수화—표준 일반화에 유리한 바람직한 요소이지만—적대적 편향에 대한 민감도 증가로 인해 강건한 일반화 오차를 악화시킬 수 있음을 드러내며, 모델 설계에서의 근본적인 상충 관계를 규명한다.

ABSTRACT

Successful deep learning models often involve training neural network architectures that contain more parameters than the number of training samples. Such overparametrized models have been extensively studied in recent years, and the virtues of overparametrization have been established from both the statistical perspective, via the double-descent phenomenon, and the computational perspective via the structural properties of the optimization landscape. Despite the remarkable success of deep learning architectures in the overparametrized regime, it is also well known that these models are highly vulnerable to small adversarial perturbations in their inputs. Even when adversarially trained, their performance on perturbed inputs (robust generalization) is considerably worse than their best attainable performance on benign inputs (standard generalization). It is thus imperative to understand how overparametrization fundamentally affects robustness. In this paper, we will provide a precise characterization of the role of overparametrization on robustness by focusing on random features regression models (two-layer neural networks with random first layer weights). We consider a regime where the sample size, the input dimension and the number of parameters grow in proportion to each other, and derive an asymptotically exact formula for the robust generalization error when the model is adversarially trained. Our developed theory reveals the nontrivial effect of overparametrization on robustness and indicates that for adversarially trained random features models, high overparametrization can hurt robust generalization.

연구 동기 및 목표

  • 과다 매개변수화가 적대적으로 훈련된 모델에서 강건한 일반화에 미치는 근본적 영향을 이해하는 것.
  • 표본 크기, 입력 차원, 모델 매개변수의 비례 증가 조건 하에서 랜덤 특징 회귀 모델의 강건한 일반화 오차를 분석하는 것.
  • 고차원 영역에서의 강건한 일반화 오차에 대해 점근적으로 정확한 공식을 유도하는 것.
  • 과다 매개변수화가 적대적 강건성에 비단순적이고 비단조화적인 역할을 할 수 있음을 드러내는 것.

제안 방법

  • 표본 크기, 입력 차원, 매개변수 수가 비례적으로 증가하는 고차원 점근적 프레임워크를 활용한다.
  • 첫 번째 레이어 가중치가 무작위로 설정된 두 층으로 구성된 신경망을 사용한다 (랜덤 특징 회귀).
  • 강건한 일반화 오차는 모레우 환경과 가우시안 등가 성질을 포함하는 변분 공식을 통해 유도된다.
  • 무작위 행렬 이론의 도구를 활용하며, 마르첸코-파스트르 분포의 스티엘제스 변환과 커널 행렬의 스펙트럼 근사치를 포함한다.
  • 한계 목표 함수의 엄격한 볼록성을 확립하여 최적 해의 유일성을 보장한다.
  • 변수 교체와 측도 집중 기법을 사용하여 최적화 지형의 점근적 행동을 분석한다.
(a) $\varepsilon=10^{-7}$
(a) $\varepsilon=10^{-7}$

실험 결과

연구 질문

  • RQ1과다 매개변수화는 적대적으로 훈련된 랜덤 특징 모델의 강건한 일반화 오차에 어떻게 영향을 미치는가?
  • RQ2과다 매개변수화된 모델에서 표준 일반화와 강건한 일반화 사이에 근본적인 상충 관계가 존재하는가?
  • RQ3고차원 영역에서 강건한 일반화 오차에 대해 점근적으로 정확한 공식을 도출할 수 있는가?
  • RQ4높은 과다 매개변수화는 표준 일반화는 향상시키지만, 강건성은 악화시키는가?

주요 결과

  • 과다 매개변수화는 일반적으로 성능 향상에 기여한다고 여겨지는 바와는 달리, 강건한 일반화 오차를 악화시킬 수 있다.
  • 모델 차원의 비례 증가 조건 하에서 강건한 일반화 오차는 점근적으로 정확한 공식으로 유도된다.
  • 분석은 과다 매개변수화가 강건성에 비단순적이고 비단조화적인 영향을 미칠 수 있음을 드러내며, 과도한 복잡성은 강건성을 해칠 수 있음을 시사한다.
  • 한계 최적화 목표 함수는 엄격한 볼록성을 가지며, 이는 강건한 일반화 및 표준 일반화 매개변수에 대해 유일한 최소화자를 보장한다.
  • 유도된 공식은 강건 오차가 표준 오차보다 높으며, 과다 매개변수화가 증가할수록 이 격차가 커짐을 보여준다.
  • 결과는 고용량임에도 불구하고 적대적으로 훈련된 모델이 여전히 취약함을 경험하는 경험적 관찰과 일치한다.
(b) $\varepsilon=0.1$
(b) $\varepsilon=0.1$

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.