Skip to main content
QUICK REVIEW

[论文解读] Symmetry-via-Duality: Invariant Neural Network Densities from Parameter-Space Correlators

Anindita Maiti, Keegan Stoner|arXiv (Cornell University)|Jun 1, 2021
Neural Networks and Applications被引用 10
一句话总结

本文提出了一种名为'symmetry-via-duality'的方法,通过在参数空间中计算的相关函数,即使在密度未知或网络不具备等变性的情况下,也能确定神经网络密度的对称性。关键贡献在于,这些相关函数在变换下的不变性揭示了输入和输出对称性,该方法在无限宽极限(恢复高斯过程结果)和有限宽训练中均得到验证,表明仅当对称性打破与真实结构对齐时,才有助于性能提升。

ABSTRACT

Parameter-space and function-space provide two different duality frames in which to study neural networks. We demonstrate that symmetries of network densities may be determined via dual computations of network correlation functions, even when the density is unknown and the network is not equivariant. Symmetry-via-duality relies on invariance properties of the correlation functions, which stem from the choice of network parameter distributions. Input and output symmetries of neural network densities are determined, which recover known Gaussian process results in the infinite width limit. The mechanism may also be utilized to determine symmetries during training, when parameters are correlated, as well as symmetries of the Neural Tangent Kernel. We demonstrate that the amount of symmetry in the initialization density affects the accuracy of networks trained on Fashion-MNIST, and that symmetry breaking helps only when it is in the direction of ground truth.

研究动机与目标

  • 开发一种在密度未知或网络不具备等变性时识别神经网络密度对称性的方法。
  • 建立神经网络参数空间与函数空间描述之间的对偶框架,用于对称性分析。
  • 证明可通过参数空间相关函数的不变性,提取网络密度的对称性。
  • 测试初始化对称性对有限宽网络泛化性能的影响,特别是在Fashion-MNIST上的表现。
  • 将该框架扩展至分析训练过程中的对称性以及神经正切核(NTK)中的对称性。

提出的方法

  • 利用参数空间与函数空间的对偶性,分析神经网络密度的对称性。
  • 使用网络输出在参数空间中的相关函数作为对称性的探测工具,依赖其在变换下的不变性。
  • 将该方法应用于检测无限宽极限(恢复高斯过程结果)和有限宽网络中的输入与输出对称性。
  • 借助格点场论类比,在离散输入集上定义函数密度,从而对非高斯过程进行严格解释。
  • 通过在Fashion-MNIST上的实验训练,测试初始化对称性对测试准确率的影响。
  • 分析训练过程中神经正切核(NTK)和参数相关性,以评估对称性的演化。
Figure 1: Test accuracy %age on Fashion-MNIST. (Left): Dependence on symmetry breaking parameters $\mu_{W}$ and $k$ for one-hot encoded labels. The error is presented in Appendix ( D ). (Right): Dependence on $\mu_{W}$ for one-cold encoded labels, showing the $95\%$ confidence interval.
Figure 1: Test accuracy %age on Fashion-MNIST. (Left): Dependence on symmetry breaking parameters $\mu_{W}$ and $k$ for one-hot encoded labels. The error is presented in Appendix ( D ). (Right): Dependence on $\mu_{W}$ for one-cold encoded labels, showing the $95\%$ confidence interval.

实验结果

研究问题

  • RQ1即使密度未知且网络不具备等变性,是否仍可确定神经网络密度的对称性?
  • RQ2参数空间相关函数的不变性特性如何揭示函数密度的输入与输出对称性?
  • RQ3在有限宽神经网络中,初始化密度的对称性在泛化中起什么作用?
  • RQ4训练过程中的对称性破缺如何影响性能,其在何种条件下具有益处?
  • RQ5该对偶框架是否可应用于分析训练过程中神经正切核和参数相关性的对称性?

主要发现

  • 即使在缺乏密度先验知识或网络等变性的情况下,仍可通过参数空间相关函数的不变性成功推断神经网络密度的对称性。
  • 在无限宽极限下,该方法成功恢复了已知的高斯过程结果,验证了其与现有理论的一致性。
  • 在有限宽网络中,初始化密度中的对称性程度直接影响Fashion-MNIST上的测试准确率,对称性越高,性能越好。
  • 训练过程中的对称性破缺仅在与数据真实对称性一致时才有助于泛化,表明归纳偏置必须具有物理意义。
  • 该框架可应用于神经正切核和训练过程中的参数相关性,表明对称性特性会动态演化,并可通过对偶性进行分析。
  • 数值实验确认,低宽条件下偏离预期对称性的现象源于有限样本波动,而非根本性对称性破缺。
Figure 2: Variation measures of $2$ -pt and $4$ -pt functions and their predicted error bounds, for $SO(3)$ and $SO(5)$ transformations of $D=3$ and $D=5$ networks, respectively.
Figure 2: Variation measures of $2$ -pt and $4$ -pt functions and their predicted error bounds, for $SO(3)$ and $SO(5)$ transformations of $D=3$ and $D=5$ networks, respectively.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。