Skip to main content
QUICK REVIEW

[论文解读] Dynamical Isometry is Achieved in Residual Networks in a Universal Way for any Activation Function

Wojciech Tarnowski, Piotr Warchoł|arXiv (Cornell University)|Sep 24, 2018
Neural Networks and Applications被引用 16
一句话总结

本文通过运用自由概率论与随机矩阵理论,推导出在初始化时输入-输出雅可比矩阵的通用谱密度公式,证明了残差神经网络(ResNets)可对任意激活函数普遍实现动态等距性。关键结果表明,通过将权重初始化方差与跳跃连接数量成反比进行调节,即可在不依赖激活函数类型的情况下实现动态等距性,从而确保不同激活函数下训练动力学的一致性。

ABSTRACT

We demonstrate that in residual neural networks (ResNets) dynamical isometry is achievable irrespectively of the activation function used. We do that by deriving, with the help of Free Probability and Random Matrix Theories, a universal formula for the spectral density of the input-output Jacobian at initialization, in the large network width and depth limit. The resulting singular value spectrum depends on a single parameter, which we calculate for a variety of popular activation functions, by analyzing the signal propagation in the artificial neural network. We corroborate our results with numerical simulations of both random matrices and ResNets applied to the CIFAR-10 classification problem. Moreover, we study the consequence of this universal behavior for the initial and late phases of the learning processes. We conclude by drawing attention to the simple fact, that initialization acts as a confounding factor between the choice of activation function and the rate of learning. We propose that in ResNets this can be resolved based on our results, by ensuring the same level of dynamical isometry at initialization.

研究动机与目标

  • 建立残差网络中动态等距性的普遍可实现性,无论使用何种激活函数。
  • 在大网络宽度与深度极限下,推导出输入-输出雅可比矩阵谱密度的通用公式。
  • 通过确保不同激活函数间具有相同的动态等距水平,消除训练动力学实验中的混杂影响。
  • 通过在CIFAR-10上训练的随机矩阵与ResNets的数值模拟,验证理论预测。
  • 提供一种有理论依据的初始化策略,使不同激活函数在学习动力学方面的比较更加公平。

提出的方法

  • 应用自由概率论与随机矩阵理论(FPT & RMT),分析深度残差网络在初始化时的信号传播。
  • 在大深度与大宽度极限下,推导出输入-输出雅可比矩阵格林函数的通用方程。
  • 定义一个单一的有效累积量参数,该参数决定奇异值谱,其值可由激活函数特性与权重方差计算得出。
  • 利用动态平均场理论,建模信号方差与梯度流经残差块的演化过程。
  • 通过引入学习到的归一化参数,将批量归一化整合到理论框架中,调整雅可比矩阵的谱统计特性。
  • 通过在随机矩阵与CIFAR-10上训练的ResNets进行数值实验,验证理论预测。

实验结果

研究问题

  • RQ1是否可以在任意激活函数下,实现残差网络的动态等距性,且其结果与函数形式无关?
  • RQ2在深度ResNets初始化时,输入-输出雅可比矩阵的何种通用谱特性起主导作用?
  • RQ3权重初始化方差与跳跃连接数量之间应满足何种关系,才能确保动态等距性?
  • RQ4初始雅可比谱在多大程度上会干扰不同激活函数在训练动力学比较中的可比性?
  • RQ5是否可以将批量归一化一致地整合到雅可比矩阵谱分析的理论框架中?

主要发现

  • 在大网络极限下,深度ResNets中输入-输出雅可比矩阵的谱密度由一个与激活函数无关的通用方程决定。
  • 当权重初始化方差与跳跃连接数量成反比时,任何激活函数在ResNets中均可实现动态等距性。
  • 决定奇异值谱的有效累积量参数,取决于激活函数、权重方差,对于某些函数还与网络深度相关。
  • 在随机矩阵与CIFAR-10训练的ResNets上的数值模拟,证实了理论预测的通用谱密度与动态等距性。
  • 当初始化确保相同动态等距水平时,不同激活函数的学习曲线更具可比性,显著降低了混杂效应。
  • 引入批量归一化后,仍保持通用谱行为,但有效累积量方程需考虑学习到的归一化参数进行修正。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。