Skip to main content
QUICK REVIEW

[論文レビュー] Benign Overfitting of Constant-Stepsize SGD for Linear Regression

Difan Zou, Jingfeng Wu|arXiv (Cornell University)|Mar 23, 2021
Stochastic Gradient Optimization Techniques参考文献 23被引用数 8
ひとこと要約

本稿は、過パラメータ化された線形回帰における定常ステップサイズのSGDに対して、反復平均またはテイル平均を用いた際の鋭い過剰リスクバインディングを提供する。これは、分散がSGD固有の有効次元に依存し、バイアスが初期反復値とデータ共分散の整合性に依存するバイアス-分散分解を明らかにする。一致する下界を確立し、テイル平均が反復平均を上回ること、およびSGDの隠れ正則化がリッジ回帰や最小ノルム補間とは根本的に異なることを示している。

ABSTRACT

There is an increasing realization that algorithmic inductive biases are central in preventing overfitting; empirically, we often see a benign overfitting phenomenon in overparameterized settings for natural learning algorithms, such as stochastic gradient descent (SGD), where little to no explicit regularization has been employed. This work considers this issue in arguably the most basic setting: constant-stepsize SGD (with iterate averaging or tail averaging) for linear regression in the overparameterized regime. Our main result provides a sharp excess risk bound, stated in terms of the full eigenspectrum of the data covariance matrix, that reveals a bias-variance decomposition characterizing when generalization is possible: (i) the variance bound is characterized in terms of an effective dimension (specific for SGD) and (ii) the bias bound provides a sharp geometric characterization in terms of the location of the initial iterate (and how it aligns with the data covariance matrix). More specifically, for SGD with iterate averaging, we demonstrate the sharpness of the established excess risk bound by proving a matching lower bound (up to constant factors). For SGD with tail averaging, we show its advantage over SGD with iterate averaging by proving a better excess risk bound together with a nearly matching lower bound. Moreover, we reflect on a number of notable differences between the algorithmic regularization afforded by (unregularized) SGD in comparison to ordinary least squares (minimum-norm interpolation) and ridge regression. Experimental results on synthetic data corroborate our theoretical findings.

研究の動機と目的

  • 定常ステップサイズのSGDの一般化行動を、明示的な正則化なしに過パラメータ化された状態で理解すること。
  • 線形回帰におけるSGDが誘導する隠れ正則化を特徴づけること、特にリッジ回帰や最小ノルム補間と比較して行うこと。
  • データ共分散行列の全固有スペクトルに依存する、過剰リスクの鋭い上界および下界を確立すること。
  • 初期重みベクトルのデータ共分散との整合性が一般化性能に与える影響を分析すること。

提案手法

  • 反復平均を用いた定常ステップサイズのSGDの過剰リスクバインディングを導出し、データ共分散の固有スペクトルに基づいてバイアスと分散に分解する。
  • 行列摂動理論とスペクトル分解を用いた新しい分析フレームワークを導入し、SGD反復値のダイナミクスを特徴づける。
  • 反復平均の一致する下界を確立し、上界の鋭さが定数要因の範囲で保証されることを証明する。
  • 最終反復に注目することでテイル平均を分析し、反復平均に比べて分散の制御が改善されることを示す。
  • 誤差をバイアスと分散の項に分解し、バイアス項がヘッセ行列の固有空間への初期重みベクトルの射影に依存することを示す。
  • トレース不等式とスペクトルバインディングを適用し、データ共分散行列の固有値に依存する非漸近的・有限標本リスクバインディングを導出する。

実験結果

リサーチクエスチョン

  • RQ1定常ステップサイズのSGDは、過パラメータ化された線形回帰でどのような条件下で良性過適合を達成するか?
  • RQ2初期重みベクトルのデータ共分散との整合性は、SGDにおける一般化にどのように影響するか?
  • RQ3線形回帰におけるSGDの反復平均とテイル平均の相対的な一般化性能はいかにか?
  • RQ4SGDの隠れ正則化は、リッジ回帰および最小ノルム補間と比べてどのように異なるか?
  • RQ5データ共分散行列の全固有スペクトルに依存する、過剰リスクの鋭い上界および下界を導出可能か?

主な発見

  • 反復平均を用いたSGDの過剰リスクは、バイアスと分散の和でバインドされ、その分散はSGDのダイナミクス固有の有効次元に依存する。
  • バイアス項は、初期重みベクトルがデータ共分散行列の固有空間への射影によって鋭く特徴づけられ、特にステップサイズと固有値に依存する。
  • 反復平均に対して一致する下界が確立され、上界の鋭さが定数要因の範囲で保証された。
  • テイル平均は反復平均よりも優れた過剰リスクバインディングを達成し、ほぼ一致する下界を持つため、一般化性能において優位性が示された。
  • SGDの隠れ正則化は、リッジ回帰および最小ノルム補間とは顕著に異なり、特に低固有値方向の扱い方において顕著である。
  • 合成データ上の実験結果は理論的発見を検証し、一般化性能が初期重みベクトルおよびステップサイズの選択に著しく依存することを示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。