Skip to main content
QUICK REVIEW

[論文レビュー] Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks

Melih Barsbey, Milad Sefidgaran|arXiv (Cornell University)|Jun 7, 2021
Stochastic Gradient Optimization Techniques参考文献 55被引用数 12
ひとこと要約

本稿は、確率的勾配降下法(SGD)における重たい尾部の重み分布と過パラメータ化されたニューラルネットワークの可逆性の間の理論的関係を確立する。SGDが大きなステップサイズ/バッチサイズ比のため重たい尾部の定常分布を示し、過パラメータ化のための混沌の伝播が生じる場合、ネットワークはモデルサイズが増加するにつれてプルーニング誤差が消えるという、明示的なℓₚ可逆性を示す。これは、単純なプルーニングがなぜ効果的であるか、そして可逆性を持つネットワークがなぜ一般化性能が良いのかを統一的に説明する。

ABSTRACT

Neural network compression techniques have become increasingly popular as they can drastically reduce the storage and computation requirements for very large networks. Recent empirical studies have illustrated that even simple pruning strategies can be surprisingly effective, and several theoretical studies have shown that compressible networks (in specific senses) should achieve a low generalization error. Yet, a theoretical characterization of the underlying cause that makes the networks amenable to such simple compression schemes is still missing. In this study, we address this fundamental question and reveal that the dynamics of the training algorithm has a key role in obtaining such compressible networks. Focusing our attention on stochastic gradient descent (SGD), our main contribution is to link compressibility to two recently established properties of SGD: (i) as the network size goes to infinity, the system can converge to a mean-field limit, where the network weights behave independently, (ii) for a large step-size/batch-size ratio, the SGD iterates can converge to a heavy-tailed stationary distribution. In the case where these two phenomena occur simultaneously, we prove that the networks are guaranteed to be '$\ell_p$-compressible', and the compression errors of different pruning techniques (magnitude, singular value, or node pruning) become arbitrarily small as the network size increases. We further prove generalization bounds adapted to our theoretical framework, which indeed confirm that the generalization error will be lower for more compressible networks. Our theory and numerical study on various neural networks show that large step-size/batch-size ratios introduce heavy-tails, which, in combination with overparametrization, result in compressibility.

研究の動機と目的

  • 過パラメータ化されたニューラルネットワークが単純なプルーニング戦略に適合する背後にあるメカニズムを特定すること。
  • 可逆性を持つネットワークがなぜ一般化性能が良いのかを説明し、過パラメータ化モデルにおける一般化の理論的理解のギャップを埋めること。
  • 深層学習における最近の2つの現象—重たい尾部のSGDダイナミクスと混沌の伝播—を、可逆性を説明する枠組みに統合すること。
  • 大きなステップサイズ/バッチサイズ比という特定のトレーニングハイパーパramータ設定下でのℓₚ可逆性に対する理論的保証を提供すること。
  • ℓₚ可逆性フレームワークに適応した一般化誤差の境界を導出し、可逆性と低一般化誤差の関係を明確にすること。

提案手法

  • 過パラメータ化されたネットワークにおけるSGDの平均場極限を分析し、ネットワークサイズが増加するにつれて重みベクトルが独立して振る舞うことを示す。
  • ステップサイズ/バッチサイズ比が大きい場合、SGDの反復が重たい尾部の定常分布に収束することを確立する。
  • 混沌の伝播(独立した重み)と重たい尾部の分布を組み合わせることで、全結合ネットワークのℓₚ可逆性を証明する。
  • 圧縮センシング理論の結果を用いて、マグニチュードプルーニング、特異値プルーニング、ノードプルーニングの各方法における圧縮誤差を評価する。
  • ℓₚ可逆性フレームワークに適応した一般化誤差の境界を導出し、可逆性と低一般化誤差の関係を結びつける。
  • 高次元球面上の濃度とエントロピーの境界を用いて、量子化された重み設定の数を制御し、誤差解析を可能にする。

実験結果

リサーチクエスチョン

  • RQ1なぜ単純なプルーニング戦略が過パラメータ化されたニューラルネットワークに対して非常に効果的なのであろうか?
  • RQ2特に重たい尾部の重み分布を示すSGDのトレーニングダイナミクスが、ネットワークの可逆性に果たす役割は何か?
  • RQ3過パラメータ化がSGDのハイパーパramータとどのように作用し、可逆性を可能にするのか?
  • RQ4特定のトレーニング条件下でℓₚ可逆性を形式的に保証できるか?
  • RQ5可逆性は一般化性能の向上を意味するのか? そして、その関係は理論的に定量化可能か?

主な発見

  • ステップサイズ/バッチサイズ比が大きい場合、SGDは重たい尾部の重み分布を生成し、これが可逆性の主要因となる。
  • 過パラメータ化された状態では、混沌の伝播によりネットワークの重みが独立して振る舞うことが保証され、効果的なプルーニングが可能になる。
  • 重たい尾部の分布と混沌の伝播が同時に発生する場合、ネットワークサイズが増加するにつれてℓₚ可逆性が明示的に保証される。
  • マグニチュードプルーニング、特異値プルーニング、ノードプルーニングの各誤差は、モデルサイズが増加するにつれて、圧縮比にかかわらず任意に小さくなる。
  • より可逆性の高いネットワークの一般化誤差は低く抑えられ、導出された一般化誤差境界が実験的観察と一致することから確認される。
  • 全結合ネットワークおよび畳み込みネットワークにおける数値実験により、理論の妥当性が検証され、重たい尾部のダイナミクスと可逆性の強い整合性が示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。