Skip to main content
QUICK REVIEW

[論文レビュー] Asymptotic Risk of Least Squares Minimum Norm Estimator under the Spike Covariance Model.

Yasaman Mahdaviyeh, Zacharie Naulet|arXiv (Cornell University)|Dec 31, 2019
Random Matrices and Applications参考文献 12被引用数 5
ひとこと要約

本稿は、パラメータ数 d が標本サイズ n よりも速やかに増加する高次元線形回帰における最小二乗最小ノルム推定量の漸近的リスクを分析する。固定された発散する固有値を有するスパイク共分散モデルの下で、リスクはゼロに収束する。これは、最小ノルム推定による補間が、高次元でも低い漸近的リスクをもたらす可能性があることを示している。

ABSTRACT

One of the recent approaches to explain good performance of neural networks has focused on their ability to fit training data perfectly (interpolate) without overfitting. It has been shown that this is not unique to neural nets, and that it happens with simpler models such as kernel regression, too Belkin et al. (2018b); Tengyuan Liang (2018). Consequently, there has been quite a few works that give conditions for low risk or optimality of interpolating models, see for example Belkin et al. (2018a, 2019b). One of the simpler models where interpolation has been studied recently is least squares solution for linear regression. In this case, interpolation is only guaranteed to happen in high dimensional setting where the number of parameters exceeds number of samples; therefore, least squares solution is not necessarily unique. However, minimum norm solution is unique, can be written in closed form, and gradient descent starting at the origin converges to it (Hastie et al., 2019). This has, at least partially, motivated several works that study risk of minimum norm least squares estimator for linear regression. Continuing in a similar vein, we study the asymptotic risk of minimum norm least squares estimator when number of parameters $d$ depends on $n$, and $\frac{d}{n} ightarrow \infty$. In this high dimensional setting, to make inference feasible, it is usually assumed that true parameters or data have some underlying low dimensional structure such as sparsity, or vanishing eigenvalues of population covariance matrix. Here, we restrict ourselves to spike covariance matrices, where a fixed finite number of eigenvalues grow with $n$ and are much larger than the rest of the eigenvalues, which are (asymptotically) in the same order. We show that in this setting the risk can vanish.

研究の動機と目的

  • d/n → ∞ である高次元線形回帰における最小ノルム最小二乗推定量のリスク行動を理解すること。
  • 過パrameter化が行われる中で、最小ノルム解による補間が低リスクをもたらす可能性があるかどうかを調査すること。
  • 有限個の発散する固有値を有するスパイク共分散モデルという構造的共分散モデル下での漸近的リスクを分析すること。
  • 最小ノルム推定量のリスクが高次元設定で消える条件を確立すること。

提案手法

  • 最小二乗最小ノルム推定量を、訓練データを適合させる制約の下でL2ノルムを最小化する制約付き最適化問題として定式化する。
  • d と n が両方とも無限大に近づく高次元漸近的設定(d/n → ∞)を仮定する。
  • 設計行列にスパイク共分散モデルを課し、固定された数の固有値が n と共に増大する一方で、残りの固有値は有界であり、漸近的に同じオーダーにとどまるようにする。
  • ランダム行列理論の道具を用いて推定量の漸近的リスクを導出する。特に、指定されたスペクトル構造下での推定量の挙動に注目する。
  • n → ∞ および d/n → ∞ の下で、リスク式の極限挙動を分析する。スパイクモデル下での解析。
  • スパイクモデル下では、リスクがゼロに収束することを確立する。これは、この設定において最小ノルム推定による補間が一貫性を持つ可能性があることを示している。

実験結果

リサーチクエスチョン

  • RQ1d/n → ∞ のとき、最小ノルム最小二乗推定量の漸近的リスクは何か?
  • RQ2最小ノルム推定量のリスクが消えるための、母集団共分散行列に関する条件は何か?
  • RQ3固定された数の発散する固有値を有するスパイク共分散構造が、補間推定量のリスク行動にどのように影響を与えるか?
  • RQ4過パrameter化が行われる高次元設定において、最小ノルム解による補間が低リスクをもたらすことができるか?
  • RQ5スパイクモデルによって示される低次元構造が、パラメータ数が標本数を上回る場合でも一貫性のある推定を可能にするか?

主な発見

  • d/n → ∞ である高次元的設定において、スパイク共分散モデル下で最小ノルム最小二乗推定量の漸近的リスクはゼロに収束する。
  • リスクの消滅はスパイク構造に起因する:固定された有限個の固有値が n と共に増大するが、残りの固有値は有界であり、漸近的に同じオーダーにとどまる。
  • この結果は、補間が保証される過パrameter化設定の中でも、設計共分散が低次元のスパイク構造を有する場合、最小ノルム解が依然として低リスクを達成できることを示している。
  • 解析により、最小ノルム解が高次元設定で inherently 高リスクを示すとは限らないことが確認された。ただし、母集団共分散がスパイク構造を有する限り、そのような状況は回避可能である。
  • この結果は、補間推定量における「穏やかな過適合(benign overfitting)」現象を支持しており、設計に構造的仮定(例:スパイク共分散)があることで低リスクが保証されることを示している。
  • この結果は、勾配降下法が過パrameter化モデルで実証的に成功する理由を理論的に裏付けるものである。なぜなら、勾配降下法は最小ノルム解に収束するが、この設定ではリスクが消えるからである。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。