Skip to main content
QUICK REVIEW

[논문 리뷰] Adversarial Learning Guarantees for Linear Hypotheses and Neural Networks

Pranjal Awasthi, Natalie S. Frank|arXiv (Cornell University)|2020. 04. 28.
Adversarial Robustness in Machine Learning참고 문헌 35인용 수 5
한 줄 요약

이 논문은 $\ell_r$-노름 편향에 대한 선형 모델과 신경망에 대해 적대적 Rademacher 복잡도에 대한 날카운 bounds를 제공하며, 이는 이전의 $\ell_\infty$-안정성 연구를 일반화한다. 일반화에 대한 차원 의존적 트레이드오프를 규명하고, 선형 분류기의 개선된 bounds를 유도하며, 대체 근사치를 사용하지 않고 단일 ReLU 유닛과 한 층의 신경망에 대한 직접적인 bounds를 제시한다.

ABSTRACT

Adversarial or test time robustness measures the susceptibility of a classifier to perturbations to the test input. While there has been a flurry of recent work on designing defenses against such perturbations, the theory of adversarial robustness is not well understood. In order to make progress on this, we focus on the problem of understanding generalization in adversarial settings, via the lens of Rademacher complexity. We give upper and lower bounds for the adversarial empirical Rademacher complexity of linear hypotheses with adversarial perturbations measured in $l_r$-norm for an arbitrary $r \geq 1$. This generalizes the recent result of [Yin et al.'19] that studies the case of $r = \infty$, and provides a finer analysis of the dependence on the input dimensionality as compared to the recent work of [Khim and Loh'19] on linear hypothesis classes. We then extend our analysis to provide Rademacher complexity lower and upper bounds for a single ReLU unit. Finally, we give adversarial Rademacher complexity bounds for feed-forward neural networks with one hidden layer. Unlike previous works we directly provide bounds on the adversarial Rademacher complexity of the given network, as opposed to a bound on a surrogate. A by-product of our analysis also leads to tighter bounds for the Rademacher complexity of linear hypotheses, for which we give a detailed analysis and present a comparison with existing bounds.

연구 동기 및 목표

  • 적대적 편향 하에서 일반화를 이해하는 데 있어 이론적 격차를 메우며, 특히 선형 모델과 신경망에 중점을 둔다.
  • 이전 연구들이 $\ell_\infty$와 같은 특정 노름에 국한되거나 대체 근사치에 의존하는 한계를 극복한다.
  • 입력 차원성과 편향 노름의 영향을 반영하는 데이터에 의존적인 날카운 적대적 Rademacher 복잡도 bounds를 제공한다.
  • Rademacher 복잡도를 통한 적대적 안정성과 일반화 사이의 직접적 연결을 수립하여 더 날카운 마진 기반 일반화 bounds를 가능하게 한다.
  • 분석 과정을 통해 비적대적 Rademacher 복잡도 bounds를 선형 분류기의 경우 개선한다.

제안 방법

  • 임의의 $\ell_r$-노름 편향($r \geq 1$) 하에서 선형 가설의 적대적 경험 Rademacher 복잡도에 대해 상한과 하한을 유도하며, 이는 이전의 $\ell_\infty$ 결과를 일반화한다.
  • 학습 샘플에 대한 분할 기반 접근법을 사용하고, Khintchine-Kahane 부등식을 적용하여 가중치 벡터 합의 $\ell_{p^*}$-노름을 근사한다.
  • Rademacher 랜덤 변수를 사용한 새로운 분석 기법을 도입하여, 차원 의존적 하한을 도출하는 적대적 셸터링을 분석한다.
  • 이중성과 노름 부등식(예: Hölder 및 Cauchy-Schwarz)을 적용하여 적대적 복잡도를 데이터 행렬의 $\ell_{p,\infty}$-노름과 연결한다.
  • 단일 ReLU 유닛에 대해 분석을 확장하여, 그들의 부호 활성화된 가중치 조합의 복잡도를 근사한다.
  • 대체 함수에 의존하지 않는 한 층의 피드포워드 신경망의 적대적 Rademacher 복잡도에 대한 직접적인 상한 bounds를 제공한다.

실험 결과

연구 질문

  • RQ1선형 모델의 적대적 Rademacher 복잡도는 입력 차원성과 편향 노름 $\ell_r$에 따라 어떻게 스케일링되는가?
  • RQ2$\ell_r$-편향 하에서 데이터 분포와 가중치 노름에 따른 적대적 일반화의 정확한 의존성은 무엇인가?
  • RQ3ReLU 유닛과 얕은 신경망에 대해 더 날카운, 데이터에 의존적인 적대적 Rademacher 복잡도 bounds를 도출할 수 있는가?
  • RQ4적대적 복잡도의 bounds는 비적대적 대비와 비교해 어떻게 다른가? 이 격차에서 차원성의 역할은 무엇인가?
  • RQ5이 분석을 통해 표준(비적대적) 선형 분류기의 일반화 bounds를 개선할 수 있는가?

주요 결과

  • 입력 차원 $d$에 대해 $\ell_r$-노름 편향 하에서 선형 모델의 적대적 Rademacher 복잡도는 추가적인 $\sqrt{d}$-의존 항을 포함하며, 이는 차원 의존적 일반화 비용을 확인한다.
  • $r \geq 1$ 인 $\ell_r$-노름에 대해, $\mathcal{O}(\tau \|X\|_{p,\infty} / \epsilon)$의 날카운 상한 bounds를 도출한다. 여기서 $\tau$는 최대 가중치 노름이고 $\epsilon$은 편향 반경이다.
  • 분석을 통해 비적대적 Rademacher 복잡도의 새로운 개선된 bounds를 도출하였으며, 이는 이전 결과보다 더 날카운 $\sqrt{d}$-의존성을 포함한다.
  • 단일 ReLU 유닛의 경우, 적대적 Rademacher 복잡도는 $\mathcal{O}(\|W^+\|_2 \|\Delta W\|_{p^*} / \epsilon)$로 bounds되며, 가중치 구조에 명시적인 의존성을 포함한다.
  • 한 층의 신경망에 대해, 대체 함수에 의존하지 않는 직접적인 상한 bounds를 제공하여 이전 연구를 향상시켰다.
  • 유도된 적대적 셸터링 bounds는 $t \leq \frac{4c_2^2(p^*)\tau^2\|X\|_{p,\infty}^2}{\epsilon^2 w_{\text{min}}^2}$를 의미하며, 여기서 $t$는 셸터링 점의 수이다. 이는 차원성과 노름을 고려한 일반화 한계를 보여준다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.