Skip to main content
QUICK REVIEW

[論文レビュー] Stationary Points of Shallow Neural Networks with Quadratic Activation Function

David Gamarnik, Eren C. Kızıldağ|arXiv (Cornell University)|Dec 3, 2019
Stochastic Gradient Optimization Techniques参考文献 62被引用数 8
ひとこと要約

この論文は、教師-生徒フレームワーク下での2層ニューラルネットワークの最適化の多様性を、2次活性化関数を用いて分析する。勾配降下法が、明示的に定量化されたエネルギー障壁の下で初期化された場合、多項式時間でグローバル最適解に収束することを確立し、過パラメータ化下でスケーリングされた単位行列によるランダム初期化が、このような初期化を確率的に満たすことを示している。ランダム行列理論を活用している。

ABSTRACT

We consider the teacher-student setting of learning shallow neural networks with quadratic activations and planted weight matrix $W^*\in\mathbb{R}^{m imes d}$, where $m$ is the width of the hidden layer and $d\le m$ is the data dimension. We study the optimization landscape associated with the empirical and the population squared risk of the problem. Under the assumption the planted weights are full-rank we obtain the following results. First, we establish that the landscape of the empirical risk admits an "energy barrier" separating rank-deficient $W$ from $W^*$: if $W$ is rank deficient, then its risk is bounded away from zero by an amount we quantify. We then couple this result by showing that, assuming number $N$ of samples grows at least like a polynomial function of $d$, all full-rank approximate stationary points of the empirical risk are nearly global optimum. These two results allow us to prove that gradient descent, when initialized below the energy barrier, approximately minimizes the empirical risk and recovers the planted weights in polynomial-time. Next, we show that initializing below this barrier is in fact easily achieved when the weights are randomly generated under relatively weak assumptions. We show that provided the network is sufficiently overparametrized, initializing with an appropriate multiple of the identity suffices to obtain a risk below the energy barrier. At a technical level, the last result is a consequence of the semicircle law for the Wishart ensemble and could be of independent interest. Finally, we study the minimizers of the empirical risk and identify a simple necessary and sufficient geometric condition on the training data under which any minimizer has necessarily zero generalization error. We show that as soon as $N\ge N^*=d(d+1)/2$, randomly generated data enjoys this geometric condition almost surely, while that ceases to be true if $N

研究の動機と目的

  • 2次活性化関数を用いた浅いニューラルネットワークの最適化の多様性を、教師-生徒設定において理解すること。
  • 勾配降下法がグローバル最適解に収束する条件を特定すること。
  • 初期化の役割が悪い局所最適解を避ける上で果たすものについて特徴づけること。
  • 一般化誤差がゼロとなるための訓練サンプルの臨界数を特定すること。
  • 最小化子がゼロ一般化誤差を達成するためのデータに関する幾何的条件を確立すること。

提案手法

  • 埋め込みられた重み行列がフルランクであると仮定したもとで、経験的リスク関数と母集団リスク関数を分析する。
  • ランク不足の重み行列と真の重み行列 $W^*$ を分離するエネルギー障壁を導入し、ランク不足の $W$ における経験的リスクの下限を定量化する。
  • 特にウィシャール分布の半円則を用いたランダム行列理論を活用し、スケーリングされた単位行列による初期化が、高確率でリスクをエネルギー障壁の下に置くことを示す。
  • サンプル数 $N$ が $d$ に対して多項式的に増加する場合、経験的リスクのすべてのフルランク近似停留点がほぼグローバル最適解であることを確立する。
  • 最小化子がゼロ一般化誤差を達成するための、訓練データに関する必要十分な幾何的条件を、$X_i X_i^T$ の生成する空間に基づいて導出する。
  • トレース不等式やスペクトルノルムの境界といった、線形代数とランダム行列理論の道具を用いて、リスクの下限を導出する。

実験結果

リサーチクエスチョン

  • RQ12次活性化関数を用いた浅いネットワークにおいて、勾配降下法がグローバル最適解に収束する条件は何か?
  • RQ2非凸最適化の多様性において、初期化が悪い局所最適解を避ける上で果たす役割は何か?
  • RQ3最小化子がゼロ一般化誤差を達成するための訓練サンプル数 $N^*$ の臨界値は何か?
  • RQ4訓練データにどのような幾何的条件が満たされると、すべての最小化子がゼロ一般化誤差を持つようになるか?
  • RQ5ランク不足の重み行列における経験的リスクはどのように振る舞い、その定量的特徴は何か?

主な発見

  • ランク不足の $W$ における経験的リスク $\widehat{\mathcal{L}}(W)$ は、明示的に定量化されたエネルギー障壁によってゼロから離れており、真の重み行列 $W^*$ と分離されている。
  • サンプル数 $N$ が $d$ に対して少なくとも多項式的に増加する場合、経験的リスク $ widehat{\mathcal{L}}(W)$ のすべてのフルランク近似停留点がほぼグローバル最適解である。
  • エネルギー障壁の下で初期化された勾配降下法は、経験的リスクを近似的に最小化し、多項式時間で $W^*$ を回復する。
  • ネットワークが十分に過パラメータ化されている場合、適切なスケーリング係数を伴う単位行列によるランダム初期化により、リスクがエネルギー障壁の下にある確率が高くなる。
  • ゼロ一般化誤差のための臨界サンプル数は $N^* = d(d+1)/2$ であり、この閾値は $N \geq N^*$ のとき、確率的に満たされる。
  • ゼロ一般化誤差を達成するための必要十分な幾何的条件は、$\{X_i X_i^T\}$ の生成する空間が、$d \times d$ 対称行列の空間と一致することであり、$N \geq d(d+1)/2$ のとき確率的に成立する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。