Skip to main content
QUICK REVIEW

[論文レビュー] Conditioning of Random Feature Matrices: Double Descent and Generalization Error

Zhijun Chen, Hayden Schaeffer|arXiv (Cornell University)|Oct 21, 2021
Sparse and Compressive Sensing Techniques参考文献 39被引用数 6
ひとこと要約

この論文は、正則化を必要とせずに、ランダム特徴行列の条件数に関する高確率的バインディングを確立し、複雑さ比 $N/m$ が $\log^{-1}(N)$ または $\log(m)$ のスケーリングとなる場合に、行列が良好に条件付けられることを示している。一般化誤差におけるダブルデセント現象を、条件数の挙動と関連づけ、$m$ および $N$ に最適にスケーリングする、次元に依存しない明示的なリスクバインディングを、最小二乗法、最小ノルム補間、スparser回帰に対して導出している。

ABSTRACT

We provide (high probability) bounds on the condition number of random feature matrices. In particular, we show that if the complexity ratio $\frac{N}{m}$ where $N$ is the number of neurons and $m$ is the number of data samples scales like $\log^{-1}(N)$ or $\log(m)$, then the random feature matrix is well-conditioned. This result holds without the need of regularization and relies on establishing various concentration bounds between dependent components of the random feature matrix. Additionally, we derive bounds on the restricted isometry constant of the random feature matrix. We prove that the risk associated with regression problems using a random feature matrix exhibits the double descent phenomenon and that this is an effect of the double descent behavior of the condition number. The risk bounds include the underparameterized setting using the least squares problem and the overparameterized setting where using either the minimum norm interpolation problem or a sparse regression problem. For the least squares or sparse regression cases, we show that the risk decreases as $m$ and $N$ increase, even in the presence of bounded or random noise. The risk bound matches the optimal scaling in the literature and the constants in our results are explicit and independent of the dimension of the data.

研究の動機と目的

  • 高次元回帰設定におけるランダム特徴行列の条件付けを理解すること。
  • 正則化なしに、ランダム特徴行列の極端な特異値および条件数に関する高確率的バインディングを確立すること。
  • 一般化誤差のダブルデセント挙動を、設計行列の条件数と関連づけること。
  • ランダム特徴を用いた最小二乗法、最小ノルム補間、スパース回帰の明示的で次元に依存しないリスクバインディングを導出すること。
  • 一般化誤差が $m$ および $N$ の増加に伴い減少することを示すこと、すなわち、有界またはランダムなノイズ下でも同様に成り立つこと。

提案手法

  • 確率的行列理論を用いて、ランダム特徴行列の従属成分の集中バインディングを導出する。
  • ランダム特徴写像のグラム行列を分析し、極端な特異値および条件数をバインドする。
  • ランダム特徴行列の制限等長定数(RIC)のバインディングを確立する。
  • 圧縮センシングからの回復の強さの結果を用いて、近似誤差をスパarsityおよび特徴の質に関連付ける。
  • 高確率的不等式を適用し、条件数が期待値から逸脱するのを制御する。
  • 条件数、回復誤差、一般化誤差をつなぐ不等式の連鎖を用いてリスクバインディングを導出する。

実験結果

リサーチクエスチョン

  • RQ1複雑さ比 $N/m$ がどのような条件下で、ランダム特徴行列が高確率的に良好に条件付けられるか?
  • RQ2ランダム特徴行列の条件数の挙動が、一般化誤差におけるダブルデセントをどのように駆動するか?
  • RQ3正則化なしに、ランダム特徴を用いた回帰のリスクバインディングを導出可能か?その最適スケーリングは何か?
  • RQ4最小二乗法、最小ノルム補間、スパース回帰のリスクバインディングは、$m$、$N$、および次元にどのように依存するか?
  • RQ5活性化関数および重み分布が、条件付けおよび一般化性能に果たす役割は何か?

主な発見

  • 正則化なしに、$N/m \asymp \log^{-1}(N)$ または $\log(m)$ の場合に、ランダム特徴行列は高確率的に良好に条件付けられる。
  • 一般化誤差におけるダブルデセントは、設計行列の条件数のダブルデセント挙動によって直接駆動される。
  • 最小二乗法およびスパース回帰のリスクバインディングは、$\mathcal{O}(N^{-1} + m^{-1/2})$ にスケーリングし、次元に依存しない明示的定数を有する。
  • 最小ノルム補間では、有界またはランダムなノイズ下でも、$m$ および $N$ の増加に伴いリスクが減少するが、正則化は不要である。
  • 条件数の下界は、$N = m$ における悪条件付けを確認し、補間閾値における一般化誤差のピークを説明している。
  • 理論的バインディングは、文献における最適スケーリングと一致しており、定数は明示的かつデータ次元 $d$ に依存しない。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。