[논문 리뷰] A proof of convergence for the gradient descent optimization method with random initializations in the training of neural networks with ReLU activation for piecewise linear target functions
이 논문은 랜덤 초기화를 가진 확률적 경사 하강법이 조각별 선형 타겟 함수를 학습하는 ReLU 신경망에서 영 위험으로 수렴함을 증명한다. 정규 초기화와 균일한 입력 분포 하에서, 저자들은 전역 최소값이 이차 미분 가능인 부분다양체를 이룬다는 것을 보여주며, 헤시안의 최대 질량 조건을 만족함으로써 비볼록 최적화의 최신 수렴 이론을 활용하여 수렴을 확립한다.
Gradient descent (GD) type optimization methods are the standard instrument to train artificial neural networks (ANNs) with rectified linear unit (ReLU) activation. Despite the great success of GD type optimization methods in numerical simulations for the training of ANNs with ReLU activation, it remains - even in the simplest situation of the plain vanilla GD optimization method with random initializations and ANNs with one hidden layer - an open problem to prove (or disprove) the conjecture that the risk of the GD optimization method converges in the training of such ANNs to zero as the width of the ANNs, the number of independent random initializations, and the number of GD steps increase to infinity. In this article we prove this conjecture in the situation where the probability distribution of the input data is equivalent to the continuous uniform distribution on a compact interval, where the probability distributions for the random initializations of the ANN parameters are standard normal distributions, and where the target function under consideration is continuous and piecewise affine linear. Roughly speaking, the key ingredients in our mathematical convergence analysis are (i) to prove that suitable sets of global minima of the risk functions are \\emph{twice continuously differentiable submanifolds of the ANN parameter spaces}, (ii) to prove that the Hessians of the risk functions on these sets of global minima satisfy an appropriate \\emph{maximal rank condition}, and, thereafter, (iii) to apply the machinery in [Fehrman, B., Gess, B., Jentzen, A., Convergence rates for the stochastic gradient descent method for non-convex objective functions. J. Mach. Learn. Res. 21(136): 1--48, 2020] to establish convergence of the GD optimization method with random initializations.
연구 동기 및 목표
- 랜덤 초기화를 가진 경사 하강법이 ReLU 신경망 학습에서 영 위험으로 수렴하는지 여부라는 열린 문제를 해결하기 위해.
- 균일한 입력 분포와 표준 정규 초기화 하에서 연속적인 조각별 애프라인 타겟 함수의 경우 수렴을 확립하기 위해.
- 위험 함수의 전역 최소값 집합이 매개변수 공간의 이차 연속 미분 가능 부분다양체임을 증명하기 위해.
- 이 부분다양체 위에서 위험 함수의 헤시안이 최대 질량 조건을 만족함을 입증하기 위해.
- 고도의 수렴 이론을 적용하여 넓이, 초기화 횟수, 반복 횟수가 증가함에 따라 거의 확실하게 영 위험으로 수렴함을 보여주기 위해.
제안 방법
- 위험 함수의 전역 최소값 집합이 신경망 매개변수 공간의 $ C^2 $-연속 부분다양체임을 증명하기 위해.
- 이 부분다양체에 제한된 위험 함수의 헤시안이 최대 질량을 가지며, 비퇴화성을 보장함을 확립하기 위해.
- 미분기하학적 도구를 사용하여 전역 최소값의 부분다양체로의 경사 흐름의 국소 수렴 행동을 분석하기 위해.
- 비볼록 최적화에서 랜덤 초기화를 고려한 Fehrman 등 (2020)의 수렴 프레임워크를 활용하기 위해.
- $ K $개의 독립적인 랜덤 초기화를 가진 확률적 경사 하강법 체계를 정의하고, 모든 실행에서의 최소 위험을 추적하기 위해.
- 집중 및 尾 꼬리 경계를 적용하여, $ K \to \infty $일 때 임의로 작은 위험을 달성할 확률이 1로 수렴함을 보여주기 위해.
실험 결과
연구 질문
- RQ1랜덤 초기화를 가진 경사 하강법이 조각별 선형 타겟 함수를 학습하는 ReLU 네트워크에서 영 위험으로 수렴하는가?
- RQ2위험 함수의 전역 최소값은 매개변수 공간에서 매끄러운 부분다양체로 특징지을 수 있는가?
- RQ3위험 함수의 헤시안은 전역 최소값 집합에서 비퇴화적(최대 질량)인가?
- RQ4랜덤 초기화 수와 반복 횟수가 증가함에 따라 수렴 결과가 거의 확실하게 성립하는가?
- RQ5주어진 기하학적 및 확률적 가정 하에서 수렴 속도를 정량화할 수 있는가?
주요 결과
- 위험 함수의 전역 최소값 집합은 신경망 매개변수 공간에서 $ C^2 $-연속 부분다양체를 이룬다.
- 이 부분다양체 위에서 위험 함수의 헤시안은 최대 질량 조건을 만족하며, 국소 안정성과 수렴을 보장한다.
- 고정된 학습률 $ \gamma \leq \mathfrak{g} = ((3N+1)(24\mathfrak{D}^5 + 16N\mathfrak{D}^7)(\sup_{x\in[a,b]} \mathfrak{p}(x)))^{-1} $ 에 대해, 랜덤 초기화 수 $ K \to \infty $ 일 때 위험은 거의 확실하게 영으로 수렴한다.
- 위험의 수렴 속도는 반복 횟수 $ n $ 에 대해 지수적이다. 초기화가 부분다양체 주변의 이웃에 있을 경우 $ \mathcal{L}(\Theta_n^{k,\gamma}) \leq \mathfrak{C} \exp(-\mathfrak{c} \gamma n) $ 를 만족한다.
- 모든 $ K $개의 독립된 실행에서의 최소 위험의 수렴 확률은 $ 1 - [\mathbb{P}(\Theta_0^{1,\gamma} \notin U)]^K $ 이하로 바운드되며, $ K \to \infty $ 일 때 1로 수렴한다.
- 결과는 컴acts한 간격에서 균일한 입력 분포, 표준 정규 초기화, 연속적인 조각별 애프라인 타겟 함수라는 가정 하에서 성립한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.