[논문 리뷰] Adam Can Converge Without Any Modification On Update Rules
이 논문은 최적화 문제를 먼저 고정하고, 그 다음에 β₁와 β₂를 튜닝하는 방식으로 기존의 접근 방식과 반대되는 새로운 이론적 프레임워크를 도입함으로써, 업데이트 규칙을 수정하지 않고도 Adam 최적화기가 임계점의 이웃으로 수렴함을 증명한다. β₂가 크고 β₁ < √β₂ < 1일 때 수렴성을 확립하며, 표준 설정인 β₁ = 0.9를 포함한 다양한 설정에서 성립한다. β₂가 증가함에 따라 발산에서 수렴으로의 단계 전이를 규명한다.
Ever since Reddi et al. 2018 pointed out the divergence issue of Adam, many new variants have been designed to obtain convergence. However, vanilla Adam remains exceptionally popular and it works well in practice. Why is there a gap between theory and practice? We point out there is a mismatch between the settings of theory and practice: Reddi et al. 2018 pick the problem after picking the hyperparameters of Adam, i.e., $(β_1, β_2)$; while practical applications often fix the problem first and then tune $(β_1, β_2)$. Due to this observation, we conjecture that the empirical convergence can be theoretically justified, only if we change the order of picking the problem and hyperparameter. In this work, we confirm this conjecture. We prove that, when $β_2$ is large and $β_1 < \sqrt{β_2}<1$, Adam converges to the neighborhood of critical points. The size of the neighborhood is propositional to the variance of stochastic gradients. Under an extra condition (strong growth condition), Adam converges to critical points. It is worth mentioning that our results cover a wide range of hyperparameters: as $β_2$ increases, our convergence result can cover any $β_1 \in [0,1)$ including $β_1=0.9$, which is the default setting in deep learning libraries. To our knowledge, this is the first result showing that Adam can converge without any modification on its update rules. Further, our analysis does not require assumptions of bounded gradients or bounded 2nd-order momentum. When $β_2$ is small, we further point out a large region of $(β_1,β_2)$ where Adam can diverge to infinity. Our divergence result considers the same setting as our convergence result, indicating a phase transition from divergence to convergence when increasing $β_2$. These positive and negative results can provide suggestions on how to tune Adam hyperparameters.
연구 동기 및 목표
- Adam 최적화기 수렴에 대한 오랜 이론-실무 격차를 해결하기 위해.
- 원본 Adam이 알고리즘 수정 없이 수렴하는 조건을 규명하기 위해.
- 이론적으로는 발산 문제를 보이고도 실무에서는 잘 작동하는 이유를 설명하기 위해.
- 실제 초파rameter 튜닝 순서와 일치하는 수렴 프레임워크를 구축하기 위해.
- β₂ 증가에 따른 발산에서 수렴으로의 행동 단계 전이를 특성화하기 위해.
제안 방법
- 최적화 문제를 먼저 고정하고, 그 다음에 β₁와 β₂를 튜닝하는 방식으로 이론적 분석을 재구성함 — 기존 연구가 초파rameter를 먼저 고정한 것과 반대.
- β₂가 크고 β₁ < √β₂ < 1일 때 Adam이 임계점의 이웃으로 수렴함을 증명.
- 강한 성장 조건 하에서, 유한한 기울기나 유한한 두 번째 순서 운동량을 요구하지 않고도 임계점으로의 수렴을 확립.
- 같은 프레임워크를 사용해 β₂가 작을 경우 (β₁, β₂) 영역의 광범위한 발산 영역을 분석.
- 수렴 이웃에서의 확률적 기울기 분산을 고려하는 새로운 분석 기법을 사용.
- β₂ 값에 따라 Adam의 수렴과 발산을 둘 다 설명하는 통합된 이론적 프레임워크를 제공.
실험 결과
연구 질문
- RQ1기존 연구에서 이론적으로 발산함에도 불구하고 실무에서는 왜 Adam이 수렴하는가?
- RQ2실제 초파rameter 설정 하에서 원본 Adam이 업데이트 규칙을 수정하지 않고도 수렴할 수 있는가?
- RQ3β₂는 Adam의 수렴 또는 발산을 결정하는 데 어떤 역할을 하는가?
- RQ4β₂ 증가에 따라 Adam의 행동에 단계 전이가 존재하는가?
- RQ5유한한 기울기나 유한한 두 번째 순서 운동량을 가정하지 않고도 수렴을 증명할 수 있는가?
주요 결과
- β₂가 크고 β₁ < √β₂ < 1일 때, Adam은 임계점의 이웃으로 수렴하며, 이 이웃의 크기는 확률적 기울기 분산에 비례한다.
- 강한 성장 조건 하에서, Adam은 임계점으로 수렴하며, 그 이웃으로의 수렴이 아니고서도 성립한다.
- 이 수렴 결과는 β₁ = 0.9를 포함한 모든 표준 β₁ 값에 적용 가능하며, 딥러닝 라이브러리의 기본 설정과 일치한다.
- 작은 β₂에 대해 발산 영역이 규명되어, β₂ 증가에 따라 발산에서 수렴으로의 단계 전이가 존재함을 시사한다.
- 이 분석은 기울기의 유한성이나 두 번째 순서 운동량의 유한성을 요구하지 않아 이전 연구보다 더 일반적인 결과를 도출한다.
- 이론적 프레임워크는 실무에서의 초파rameter 튜닝 순서와 일치하여, 이론과 실무 간의 괴리 문제를 해결한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.