Skip to main content
QUICK REVIEW

[论文解读] No Spurious Local Minima in Deep Quadratic Networks

Abbas Kazemipour, Brett W. Larsen|arXiv (Cornell University)|Dec 31, 2019
Stochastic Gradient Optimization Techniques参考文献 3被引用 6
一句话总结

本文证明了当神经元数量不少于输入维度且使用训练样本范数作为回归器时,具有二次激活函数的深层过参数化神经网络在其二次损失景观中不存在虚假局部极小值。理论分析表明所有局部极小值均为全局极小值,实证结果证实了向全局极小值的收敛,建立了与数据分布无关的有利优化特性。

ABSTRACT

Despite their practical success, a theoretical understanding of the loss landscape of neural networks has proven challenging due to the high-dimensional, non-convex, and highly nonlinear structure of such models. In this paper, we characterize the training landscape of the quadratic loss landscape for neural networks with quadratic activation functions. We prove existence of spurious local minima and saddle points which can be escaped easily with probability one when the number of neurons is greater than or equal to the input dimension and the norm of the training samples is used as a regressor. We prove that deep overparameterized neural networks with quadratic activations benefit from similar nice landscape properties. Our theoretical results are independent of data distribution and fill the existing gap in theory for two-layer quadratic neural networks. Finally, we empirically demonstrate convergence to a global minimum for these problems.

研究动机与目标

  • 理论表征具有二次激活函数的深层神经网络的损失景观。
  • 研究此类网络在训练动态中是否存在虚假局部极小值。
  • 建立所有局部极小值均为全局极小值的条件,以确保优化收敛。
  • 证明这些有利的景观特性与数据分布无关。
  • 通过实证验证在二次神经网络中可观察到向全局极小值的收敛。

提出的方法

  • 使用二次激活函数分析深层神经网络的二次损失景观。
  • 证明当神经元数量大于或等于输入维度时,虚假局部极小值不存在。
  • 使用训练样本的范数作为回归器,以确保有利的优化特性。
  • 在指定的过参数化条件下,建立所有局部极小值均为全局极小值。
  • 应用与数据分布无关的理论分析,以推广结果至不同数据设置。
  • 通过实证实验验证理论发现,显示向全局极小值的收敛。

实验结果

研究问题

  • RQ1深层二次神经网络的训练景观中是否存在虚假局部极小值?
  • RQ2在过参数化的二次网络中,虚假局部极小值在何种条件下可被避免?
  • RQ3能否独立于数据分布证明虚假局部极小值的不存在性?
  • RQ4将训练样本的范数用作回归器如何影响损失景观?
  • RQ5在理论条件下,能否在二次神经网络中实证观察到全局收敛?

主要发现

  • 当神经元数量不少于输入维度时,深层过参数化二次神经网络中不存在虚假局部极小值。
  • 在指定的过参数化和回归器条件下,二次损失景观中的所有局部极小值均为全局极小值。
  • 理论结果与底层数据分布无关,增强了泛化能力。
  • 鞍点存在,但可以概率为一逃离,支持高效的优化。
  • 实证结果证实了在该类网络训练中可收敛至全局极小值。
  • 由于二次激活函数的结构和过参数化,有利的景观特性得以保持。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。