Skip to main content
QUICK REVIEW

[论文解读] High-dimensional limit theorems for SGD: Effective dynamics and critical scaling

Gérard Ben Arous, Reza Gheissari|arXiv (Cornell University)|Jun 8, 2022
Markov Chains and Monte Carlo Methods被引用 7
一句话总结

本文建立了常数学习率下随机梯度下降(SGD)的高维极限定理,表明SGD的统计摘要收敛到由常微分方程(ODE)或随机微分方程(SDE)所支配的有效动力系统,具体取决于初始化方式与学习率缩放。识别出一个关键的缩放区间,在此区间内梯度流被新项修正,导致非平凡的相图以及以正概率出现的次优收敛——该现象可通过过参数化得到缓解。

ABSTRACT

We study the scaling limits of stochastic gradient descent (SGD) with constant step-size in the high-dimensional regime. We prove limit theorems for the trajectories of summary statistics (i.e., finite-dimensional functions) of SGD as the dimension goes to infinity. Our approach allows one to choose the summary statistics that are tracked, the initialization, and the step-size. It yields both ballistic (ODE) and diffusive (SDE) limits, with the limit depending dramatically on the former choices. We show a critical scaling regime for the step-size, below which the effective ballistic dynamics matches gradient flow for the population loss, but at which, a new correction term appears which changes the phase diagram. About the fixed points of this effective dynamics, the corresponding diffusive limits can be quite complex and even degenerate. We demonstrate our approach on popular examples including estimation for spiked matrix and tensor models and classification via two-layer networks for binary and XOR-type Gaussian mixture models. These examples exhibit surprising phenomena including multimodal timescales to convergence as well as convergence to sub-optimal solutions with probability bounded away from zero from random (e.g., Gaussian) initializations. At the same time, we demonstrate the benefit of overparametrization by showing that the latter probability goes to zero as the second layer width grows.

研究动机与目标

  • 开发常数学习率下高维设置中SGD的统一尺度极限框架。
  • 刻画初始化、学习率缩放及参数区域选择如何影响统计摘要的有效动力系统。
  • 识别一个关键缩放区间,在此区间内梯度流被新项修正,从而改变收敛行为。
  • 分析过参数化对收敛至次优解的影响。
  • 在带信号的矩阵/张量模型及高斯混合分类的两层网络上验证理论。

提出的方法

  • 作者在维数趋于无穷时,对SGD的有限维统计摘要推导出极限定理,且仅需较弱的正则性条件。
  • 根据初始化与学习率缩放,建立其收敛至ODE(弹道相)或SDE(扩散相)的结论。
  • 通过统计摘要轨迹的泛函中心极限定理推导出有效动力系统。
  • 识别出学习率缩放为 $ O(1/ ext{dimension}) $ 的关键缩放区间,导致ODE极限中出现校正项。
  • 该方法可追踪特定统计量,如损失、权重幅值以及与真实值的相关性。
  • 分析利用了高维设置下的高斯近似与矩界,包括ReLU与Sigmoid激活函数的情形。

实验结果

研究问题

  • RQ1在常数学习率下,SGD统计摘要的轨迹在高维极限中如何收敛?
  • RQ2是什么决定了有效动力系统为弹道型(ODE)还是扩散型(SDE)?
  • RQ3学习率缩放在决定有效动力系统与收敛行为中起什么作用?
  • RQ4随机初始化在非凸问题中如何影响收敛至次优解?
  • RQ5过参数化能否消除收敛至次优解的概率?

主要发现

  • SGD统计摘要的有效动力系统根据初始化与学习率缩放收敛至ODE或SDE,且在关键缩放区间内,ODE极限中出现新校正项。
  • 在关键缩放区间内,ODE极限与总体损失的梯度流一致,但包含一项校正,改变相图并可能导致收敛至次优不动点。
  • 对于XOR型高斯混合模型的两层网络,SGD在随机初始化下以正概率收敛至次优解。
  • 第二层的过参数化可降低收敛至次优解的概率,且随着第二层宽度增加,该概率趋于零。
  • 该方法适用于带信号的矩阵与张量模型,揭示了多时间尺度的收敛行为及复杂的扩散极限。
  • 分析表明,扩散极限可能退化,且收敛行为对统计摘要与初始化的选择极为敏感。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。