[논문 리뷰] Rethinking the limiting dynamics of SGD: modified loss, phase space oscillations, and anomalous diffusion
이 논문은 확률적 경사하강법(SGD)이 수렴 후에도 매개변수 공간에서 비정상 확산(anomalous diffusion)를 통해 깊은 신경망을 계속 구동시킨다는 것을 밝혀냈다. 이 경우 이동 거리는 반복 횟수의 거듭제곱 법칙에 따라 증가한다. 연속 시간의 과소된 랑주뱅 모델과 선형 회귀에서의 포커-플랑크 분석을 통해, 기울기 노이즈와 헤시안 행렬에 의해 결정되는 수정된 손실 함수와 확률 흐름이 위상 공간 진동과 비자명한 확산 지수를 이끄는 핵심 요소임을 규명하였다.
In this work we explore the limiting dynamics of deep neural networks trained with stochastic gradient descent (SGD). We find empirically that long after performance has converged, networks continue to move through parameter space by a process of anomalous diffusion in which distance travelled grows as a power law in the number of gradient updates with a nontrivial exponent. We reveal an intricate interaction between the hyperparameters of optimization, the structure in the gradient noise, and the Hessian matrix at the end of training that explains this anomalous diffusion. To build this understanding, we first derive a continuous-time model for SGD with finite learning rates and batch sizes as an underdamped Langevin equation. We study this equation in the setting of linear regression, where we can derive exact, analytic expressions for the phase space dynamics of the parameters and their instantaneous velocities from initialization to stationarity. Using the Fokker-Planck equation, we show that the key ingredient driving these dynamics is not the original training loss, but rather the combination of a modified loss, which implicitly regularizes the velocity, and probability currents, which cause oscillations in phase space. We identify qualitative and quantitative predictions of this theory in the dynamics of a ResNet-18 model trained on ImageNet. Through the lens of statistical physics, we uncover a mechanistic origin for the anomalous limiting dynamics of deep neural networks trained with SGD.
연구 동기 및 목표
- 학습 손실이 안정화된 후에도 지속되는 비평형 동역학을 신경망 매개변수에서 이해한다.
- SGD 최적화의 한계 단계에서 매개변수 공간의 비정상 확산을 이끄는 메커니즘을 규명한다.
- 초기화 조정, 기울기 노이즈 구조, 헤시안 행렬이 장기적 동역학을 어떻게 함께 형성하는지 드러낸다.
- 유한 배치 크기의 SGD를 연속 시간 모델로 기술하여 속도에 의존하는 동역학과 위상 공간 진동을 포괄하는 모델을 수립한다.
- 이론적 예측을 ImageNet에서 훈련된 ResNet-18에서 검증하여 통계역학과 딥러닝 동역학을 연결한다.
제안 방법
- 유한 학습률과 배치 크기를 고려한 연속 시간 과소된 랑주뱅 방정식을 수립하여 유한 배치 크기의 SGD를 모델링한다.
- 포커-플랑크 방정식을 사용하여 선형 회귀에서 매개변수 및 속도 동역학에 대한 정확한 해석적 해를 유도한다.
- 원래의 학습 손실과는 다름없이 매개변수 공간에서 속도를 암묵적으로 정규화하는 수정된 손실 함수를 도입한다.
- 위상 공간에서의 확률 흐름을 분석하여 매개변수 궤적의 지속적인 진동 행동을 설명한다.
- 기울기 노이즈 구조, 헤시안 곡률, 그리고 그로 인한 비자명한 확산 지수 간의 상호작용을 특성화한다.
- ImageNet에서 훈련된 ResNet-18의 위상 공간 동역학을 분석하여 이론적 예측을 검증한다.
실험 결과
연구 질문
- RQ1학습 손실이 안정화된 후에도 신경망 매개변수가 지속적으로 수렴하지 않는 움직임을 이끄는 요소는 무엇인가?
- RQ2기울기 노이즈, 헤시안 곡률, 최적화 하이퍼파rameter 간의 상호작용이 장기적 매개변수 동역학을 어떻게 형성하는가?
- RQ3속도 정규화와 확률 흐름은 매개변수-속도 위상 공간에서의 진동 행동을 어떻게 생성하는가?
- RQ4매개변수 공간에서의 비정상 확산 지수는 수정된 손실 함수와 포커-플랑크 동역학으로 어느 정도 설명될 수 있는가?
- RQ5선형 모델에서 유도된 이론적 프레임워크는 ResNet-18과 같은 깊은 비선형 네트워크의 동역학을 설명하는 데 확장 가능한가?
주요 결과
- 수렴 후에도 깊은 신경망은 매개변수 공간에서 비정상 확산을 계속 보이며, 이동 거리는 반복 횟수의 거듭제곱 법칙에 따라 증가한다.
- 확산 지수는 비자명하며, 헤시안 행렬, 기울기 노이즈 구조, 최적화 하이퍼파rameter 간의 상호작용에 의해 결정된다.
- 동역학의 핵심 원동력은 원래의 학습 손실이 아니라, 매개변수 공간에서 속도를 암묵적으로 정규화하는 수정된 손실 함수이다.
- 위상 공간의 확률 흐름은 지속적인 진동을 유도하여 시스템이 단순한 평형 상태에 도달하는 것을 방지한다.
- 과소된 랑주뱅 동역학을 갖는 선형 회귀 모델에서 유도된 이론적 예측은 ImageNet에서 훈련된 ResNet-18에서 정량적으로 검증되었다.
- 한계 동역학은 속도 정규화와 비평형 흐름 간의 균형에 의해 지배되며, 이는 통계역학 원리에 뿌리를 두고 있다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.