Skip to main content
QUICK REVIEW

[論文レビュー] When Does Preconditioning Help or Hurt Generalization?

Шун-ичи Амари, Jimmy Ba|arXiv (Cornell University)|Jun 18, 2020
Sparse and Compressive Sensing Techniques参考文献 88被引用数 11
ひとこと要約

本稿は、正則化なしの過パラメータ化線形回帰における前処理付き勾配降下法の下で、汎化誤差の正確な漸近的バイアス・バリアンス分解を提供しており、ラベルにノイズがある、モデルが誤って指定されている、または信号が特徴量とずれている状況では、自然勾配降下法(NGD)が勾配降下法(GD)よりも一般化性能が優れていることを示している。逆に、ラベルがクリアで、モデルが適切に指定されており、信号が特徴量と整合している場合には、GDがNGDよりも一般化性能が優れている。

ABSTRACT

While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a more nuanced view on how the extit{implicit bias} of first- and second-order methods affects the comparison of generalization properties. We provide an exact asymptotic bias-variance decomposition of the generalization error of overparameterized ridgeless regression under a general class of preconditioner $\boldsymbol{P}$, and consider the inverse population Fisher information matrix (used in NGD) as a particular example. We determine the optimal $\boldsymbol{P}$ for both the bias and variance, and find that the relative generalization performance of different optimizers depends on the label noise and the "shape" of the signal (true parameters): when the labels are noisy, the model is misspecified, or the signal is misaligned with the features, NGD can achieve lower risk; conversely, GD generalizes better than NGD under clean labels, a well-specified model, or aligned signal. Based on this analysis, we discuss several approaches to manage the bias-variance tradeoff, and the potential benefit of interpolating between GD and NGD. We then extend our analysis to regression in the reproducing kernel Hilbert space and demonstrate that preconditioned GD can decrease the population risk faster than GD. Lastly, we empirically compare the generalization error of first- and second-order optimizers in neural network experiments, and observe robust trends matching our theoretical analysis.

研究の動機と目的

  • 第一および第二順位の最適化法のimplicit biasが一般化性能に与える影響を理解すること。
  • 前処理が過パラメータ化されたリッジレス回帰における一般化性能に与える影響を調査すること。
  • NGDがGDを上回る、および逆にGDがNGDを上回る条件を同定すること。
  • 理論的知見を再生核ヒルベルト空間へと拡張し、ニューラルネットワーク実験で検証すること。

提案手法

  • 過パラメータ化されたリッジレス回帰において、ランダム行列理論を用いて一般化誤差の正確な漸近的バイアス・バリアンス分解を導出する。
  • 一般クラスの前処理行列Pを検討し、NGDで用いられる母集団フィッシャー情報行列の逆行列(P⁻¹)を特別なケースとして含む。
  • 時間に依存しない前処理行列のもとで、母集団リスクをバイアスおよびバリアンス成分の形で計算する。
  • バイアスとバリアンスを別々に最小化する最適なPを特定し、一般化性能におけるトレードオフを明らかにする。
  • 再帰的核ヒルベルト空間における回帰への分析を拡張し、前処理付きGDがGDよりも母集団リスクを速やかに減少させることを示す。
  • ニューラルネットワークにおける第一および第二順位最適化法の一般化誤差を実験的に評価し、理論的傾向を検証する。

実験結果

リサーチクエスチョン

  • RQ1NGDによる前処理がGDと比較して一般化性能を向上させる条件は何か?
  • RQ2ラベルノイズやモデルの誤って指定された状況が、NGDとGDの相対的な一般化性能に与える影響は何か?
  • RQ3特徴量と信号がずれている場合、前処理付き最適化法のバイアス・バリアンストレードオフに果たす役割は何か?
  • RQ4GDとNGDの間を補間することで、バイアスとバリアンスをバランスさせ、一般化性能を向上させられるか?
  • RQ5核法における前処理が収束速度および母集団リスクに与える影響は何か?

主な発見

  • ラベルにノイズがある、モデルが誤って指定されている、または信号が特徴量とずれている場合には、NGDがGDよりも一般化性能が優れている。
  • ラベルがクリアで、モデルが適切に指定されており、信号が特徴量と整合している場合には、GDがNGDよりも一般化性能が優れている。
  • バイアスとバリアンスの最適な前処理行列は異なり、一般化性能における根本的なトレードオフが存在する。
  • 特別なケースγ=2では、GDとNGDのバイアスが等しくなる臨界指数r*は正確に-1/2である。
  • 再帰的核ヒルベルト空間設定において、前処理付きGDはGDよりも母集団リスクを速やかに減少させる。
  • ニューラルネットワーク実験では、理論的バイアス・バリアンス分析と整合する安定した経験的傾向が観察された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。