Skip to main content
QUICK REVIEW

[论文解读] Comparing Dynamics: Deep Neural Networks versus Glassy Systems

Marco Baity‐Jesi, Levent Sagun|Mar 19, 2018
Time Series Analysis and Forecasting被引用 20
一句话总结

本文利用统计物理方法,将过参数化的深度神经网络(DNNs)的训练动力学与玻璃态系统进行比较。研究发现,由于平坦方向数量的增加,DNNs 在损失函数景观底部表现出缓慢的扩散动力学,且无势垒穿越——尽管与平均场玻璃态系统在衰老行为上存在一些相似性,但这种差异使其区别于后者。

ABSTRACT

We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are (1) the complexity of the loss landscape and of the dynamics within it, and (2) to what extent DNNs share similarities with glassy systems. Our findings, obtained for different architectures and datasets, suggest that during the training process the dynamics slows down because of an increasingly large number of flat directions. At large times, when the loss is approaching zero, the system diffuses at the bottom of the landscape. Despite some similarities with the dynamics of mean-field glassy systems, in particular, the absence of barrier crossing, we find distinctive dynamical behaviors in the two cases, showing that the statistical properties of the corresponding loss and energy landscapes are different. In contrast, when the network is under-parametrized we observe a typical glassy behavior, thus suggesting the existence of different phases depending on whether the network is under-parametrized or over-parametrized.

研究动机与目标

  • 探究深度神经网络(DNNs)是否在损失函数景观探索中表现出与玻璃态系统相似的动力学特性。
  • 确定过参数化如何改变损失函数景观的统计特性,相较于参数不足的网络。
  • 利用离散系统统计物理方法,分析DNN训练过程中是否存在势垒穿越。
  • 识别DNN中的衰老动力学与缓慢松弛现象是否源于类玻璃态行为,或源于与过参数化相关的独特机制。
  • 基于网络容量与景观结构,探索是否存在从易学习与难学习阶段之间的相变。

提出的方法

  • 使用随机梯度下降(SGD)在多种架构和数据集上对DNNs的训练动力学进行数值分析。
  • 应用玻璃态系统统计物理方法,包括时间相关函数与均方位移的分析。
  • 使用时间依赖相关函数 $\Delta(t_w, t_w + t)$ 检测衰老行为,并区分玻璃态与非玻璃态动力学。
  • 通过减少模型A中的神经元数量,比较过参数化与参数不足网络的损失景观特性。
  • 通过小时间尺度下均方位移的坍缩,研究爱德华-安德森参数与局部极小值陷阱。
  • 通过分析大 $t_w$ 时相关函数 $\Delta(t_w, t_w + t)$ 的形状,评估DNNs与平均场玻璃态系统之间的定性差异。

实验结果

研究问题

  • RQ1过参数化DNNs的训练动力学在多大程度上类似于平均场玻璃态系统?
  • RQ2DNNs中缺乏势垒穿越是否表明其损失函数景观的统计结构与玻璃态系统存在根本性差异?
  • RQ3过参数化如何影响平坦方向的出现以及由此产生的缓慢动力学?
  • RQ4DNNs中是否存在从'易学习'阶段(过参数化)到'难学习'阶段(参数不足)的相变?
  • RQ5DNNs的动力学行为是否可由宽广平坦的吸引盆而非局部极小值的玻璃态陷阱来解释?

主要发现

  • 在训练过程中,由于平坦方向数量的增加,DNNs 在损失函数景观底部表现出缓慢的扩散动力学。
  • 势垒穿越在DNN训练中不发挥显著作用,与系统未被深局部极小值捕获的观察一致。
  • 在过参数化网络中,系统达到接近零损失,并表现出类似衰老的动力学,但其统计特性与平均场玻璃态系统存在显著差异。
  • 当网络参数不足时,均方位移在小时间尺度下出现坍缩,且随 $t_w$ 增加,表明系统被陷阱捕获并表现出玻璃态衰老行为。
  • 在参数不足的网络中,损失函数无法达到零,而是渐近趋于较高值,表明泛化能力差或收敛缓慢。
  • 作者推测存在一个相变:从'易学习'相(过参数化,平坦景观,快速学习)到'难学习'相(参数不足,粗糙景观,玻璃态动力学),类似于组合优化中的相变。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。