Skip to main content
QUICK REVIEW

[논문 리뷰] The Benefits of Over-parameterization at Initialization in Deep ReLU Networks

Devansh Arpit, Yoshua Bengio|arXiv (Cornell University)|2019. 01. 11.
Stochastic Gradient Optimization Techniques참고 문헌 23인용 수 14
한 줄 요약

이 논문은 히 초기화된 깊은 ReLU 신경망이 초깃값에서 노름 유지 성질을 보임을 보여준다: 은닉 활성화의 노름은 입력 노름과 같고, 가중치 기울기의 노름은 입력 노름과 출력 오차 노름의 곱과 같다. 이러한 성질은 유한한 너비를 가진 네트워크에서 유도된 유한한 하한 너비 조건 하에 성립하며, 이는 이전 연구에서 요구된 무한한 너비 또는 i.i.d. 입력 데이터 분포와 같은 가정을 완화한다. 이 성질들은 유한한 데이터셋에 대해 높은 확률로 PAC 분석을 통해 입증된다.

ABSTRACT

It has been noted in existing literature that over-parameterization in ReLU networks generally improves performance. While there could be several factors involved behind this, we prove some desirable theoretical properties at initialization which may be enjoyed by ReLU networks. Specifically, it is known that He initialization in deep ReLU networks asymptotically preserves variance of activations in the forward pass and variance of gradients in the backward pass for infinitely wide networks, thus preserving the flow of information in both directions. Our paper goes beyond these results and shows novel properties that hold under He initialization: i) the norm of hidden activation of each layer is equal to the norm of the input, and, ii) the norm of weight gradient of each layer is equal to the product of norm of the input vector and the error at output layer. These results are derived using the PAC analysis framework, and hold true for finitely sized datasets such that the width of the ReLU network only needs to be larger than a certain finite lower bound. As we show, this lower bound depends on the depth of the network and the number of samples, and by the virtue of being a lower bound, over-parameterized ReLU networks are endowed with these desirable properties. For the aforementioned hidden activation norm property under He initialization, we further extend our theory and show that this property holds for a finite width network even when the number of data samples is infinite. Thus we overcome several limitations of existing papers, and show new properties of deep ReLU networks at initialization.

연구 동기 및 목표

  • 초기화 시 깊은 ReLU 신경망의 새로운 이론적 성질을 규명하고 형식화하여 훈련 동역학 향상에 기여한다.
  • 이전 연구에서 흔히 요구되던 무한한 너비 또는 i.i.d. 입력 데이터 분포 가정을 완화하기 위해, 유한 너비 보장을 도출한다.
  • PAC 분석을 통해 활성화 및 기울기 노름 동등성 성질이 유한한 데이터셋에 대해서도 높은 확률로 성립함을 입증한다.
  • 이 유익한 성질들이 점점 커지는 근사값에서만 나타나는 것이 아니라, 유도된 너비 하한 조건 하에 유한 너비 네트워크에서도 성립함을 보여준다.

제안 방법

  • PAC 분석을 사용하여 히 초기화된 깊은 ReLU 신경망에서 활성화 및 기울기 노름에 대한 높은 확률 경계를 도출한다.
  • 활성화 및 기울기 노름 동등성 성질이 높은 확률로 성립하도록 하는 유한한 너비 하한을 유도한다.
  • 집중 불등식(예: 허프딩 유형 경계)을 적용하여, 모든 층에서 은닉 활성화의 노름이 입력 노름에 가까이 유지됨을 보여준다.
  • 유니온 바운드와 벡터 투영 레마를 활용하여 노름 유지 성질을 전체 네트워크 깊이로 확장한다.
  • 가중치 행렬에서 i.i.d. 정규분포 및 베르누이 랜덤 변수의 성질을 활용하여, 초기화 시 ReLU 신경망의 행동을 모델링한다.
  • 전방 및 역방향 전파 분석을 통합하여 기울기 노름이 입력 노름과 출력 오차 노름의 곱에 비례함을 증명한다.

실험 결과

연구 질문

  • RQ1유한 너비 네트워크에서 히 초기화된 깊은 ReLU 신경망이 층 간 활성화 노름을 유지하는가?
  • RQ2초기화 시 깊은 ReLU 신경망에서 기울기 노름이 입력 노름과 출력 오차 노름의 곱에 비례함을 입증할 수 있는가?
  • RQ3이러한 노름 유지 성질이 높은 확률로 성립하기 위한 최소 유한 너비는 얼마인가?
  • RQ4무한한 너비 또는 i.i.d. 입력 데이터 분포를 가정하지 않고 이러한 성질들이 어떻게 유지되는가?

주요 결과

  • 각 층의 은닉 활성화 노름은 높은 확률로 입력 노름과 동일하다. 즉, \|\mathbf{h}^l\|_2 \approx \|\mathbf{x}\|_2 for all l.
  • 각 층의 가중치 기울기의 프로베니우스 노름은 입력 노름과 출력 오차 노름의 곱과 약간 다르지 않다. 즉, \|\partial \ell / \partial \mathbf{W}^l\|_F \approx \|\delta(\mathbf{x},\mathbf{y})\|_2 \cdot \|\mathbf{x}\|_2.
  • 이러한 노름 동등성 성질이 높은 확률로 성립하도록 하는 유한한 너비 하한이 존재하며, 깊이와 샘플 수에 따라 달라진다.
  • 활성화 노름 동등성 성질은 너비가 유도된 유한한 하한 너비를 초과할 경우, 무한한 데이터 스트림이 존재하더라도 유지된다.
  • 이론적 경계는 i.i.d. 입력 데이터나 무한한 너비를 가정하지 않고 도출되었으며, 이는 이전의 점근적 결과보다 향상된 것이다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.