Skip to main content
QUICK REVIEW

[论文解读] Neural Tangent Kernel Beyond the Infinite-Width Limit: Effects of Depth and Initialization

Mariia Seleznova, Gitta Kutyniok|arXiv (Cornell University)|Feb 1, 2022
Neural Networks and Applications被引用 5
一句话总结

本文研究了深度全连接ReLU网络在无限宽极限之外的神经正切核(NTK),表明在混沌相和临界相中,由于依赖于初始化的动力学,NTK的变异性随深度呈指数增长;而在有序相中,NTK保持稳定。作者推导了Poole等人(2016)所识别的三种相——有序相、混沌相和临界相(EOC)——在无限深度与无限宽度极限下的NTK方差的精确表达式,揭示了只有在有序相中,训练过程中NTK才能保持恒定。

ABSTRACT

Neural Tangent Kernel (NTK) is widely used to analyze overparametrized neural networks due to the famous result by Jacot et al. (2018): in the infinite-width limit, the NTK is deterministic and constant during training. However, this result cannot explain the behavior of deep networks, since it generally does not hold if depth and width tend to infinity simultaneously. In this paper, we study the NTK of fully-connected ReLU networks with depth comparable to width. We prove that the NTK properties depend significantly on the depth-to-width ratio and the distribution of parameters at initialization. In fact, our results indicate the importance of the three phases in the hyperparameter space identified in Poole et al. (2016): ordered, chaotic and the edge of chaos (EOC). We derive exact expressions for the NTK dispersion in the infinite-depth-and-width limit in all three phases and conclude that the NTK variability grows exponentially with depth at the EOC and in the chaotic phase but not in the ordered phase. We also show that the NTK of deep networks may stay constant during training only in the ordered phase and discuss how the structure of the NTK matrix changes during training.

研究动机与目标

  • 理解深度和初始化如何影响无限宽极限之外的神经正切核(NTK)。
  • 分析NTK的统计特性,特别是其在深度与宽度相近的深层网络中训练过程中的方差和恒定性。
  • 确定在何种初始化条件下,NTK在训练过程中保持恒定,尤其是在Poole等人(2016)所识别的三种相——有序相、混沌相和临界相(EOC)——的背景下。
  • 推导三种相在无限深度与无限宽度极限下NTK方差的精确表达式。

提出的方法

  • 在联合无限深度与无限宽度极限下,对全连接ReLU网络中的NTK进行理论分析。
  • 推导三种相——有序相、混沌相和临界相——在初始化时的NTK方差(二阶矩与一阶矩平方的比值)的精确表达式。
  • 利用随机矩阵理论和递推矩传播方法,刻画NTK统计量随网络深度的演化。
  • 应用自助抽样法估计经验NTK方差的标准误,确保理论预测的可靠统计验证。
  • 将理论NTK方差与蒙特卡洛模拟在不同深度与宽度比及初始化参数下的经验结果进行比较。
  • 基于梯度范数和参数传播在各相中的特定行为,识别NTK在训练过程中保持恒定的条件。

实验结果

研究问题

  • RQ1当深度和宽度同时趋于无穷大时,深度全连接ReLU网络中的神经正切核(NTK)行为如何?
  • RQ2深度与宽度比以及初始化超参数在决定NTK训练过程中的变异性与恒定性方面起什么作用?
  • RQ3在三种相——有序相、混沌相或临界相——中,NTK在训练过程中保持恒定的是哪一种?原因是什么?
  • RQ4在三种相中,NTK在初始化时的方差如何随深度变化?
  • RQ5无限宽极限在多大程度上无法捕捉深层网络的真实行为,特别是在NTK稳定性方面?

主要发现

  • 在混沌相和临界相(EOC)中,NTK方差随深度呈指数增长,表明训练过程中NTK具有高度变异性。
  • 在有序相中,由于梯度范数随深度减小,NTK在训练过程中保持恒定,从而实现稳定。
  • 在无限深度与无限宽度极限下,NTK方差的解析表达式被推导出来,且其表现具有相依赖性,三种相中表现出不同的标度行为。
  • 经验结果证实,NTK方差与理论预测一致,自助法估计的连续误差条验证了研究结果的稳健性。
  • 深层网络的NTK仅在有序相中训练过程中保持恒定;而在混沌相和EOC相中,由于参数敏感性增加,NTK显著演化。
  • 在混沌相和EOC相中,NTK矩阵的结构在训练过程中发生显著变化,非对角元素相对于对角元素增长,表明样本间存在强烈的相互作用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。