Skip to main content
QUICK REVIEW

[論文レビュー] Overfitting Can Be Harmless for Basis Pursuit, But Only to a Degree

Peizhong Ju, Xiaojun Lin|arXiv (Cornell University)|Feb 2, 2020
Sparse and Compressive Sensing Techniques参考文献 39被引用数 8
ひとこと要約

本稿は、$p$(特徴量)が $n$(サンプル数)を上回る過パラメータ化された状況において、$\ell_1$-ノルム最小化手法である Basis Pursuit (BP) を分析している。過学習によるBPの性能劣化が無害であることが示され、一般化誤差は $p$ が $n$ の指数関数的スケールに達するまで減少することが判明した。主な結果は、$p$ とともに減少する非漸近的かつ高確率的なモデル誤差の上界であり、これは最小 $\ell_2$-ノルム解とは異なり、二重降下(double-descent)の挙動を示すものである。

ABSTRACT

Recently, there have been significant interests in studying the so-called "double-descent" of the generalization error of linear regression models under the overparameterized and overfitting regime, with the hope that such analysis may provide the first step towards understanding why overparameterized deep neural networks (DNN) still generalize well. However, to date most of these studies focused on the min $\ell_2$-norm solution that overfits the data. In contrast, in this paper we study the overfitting solution that minimizes the $\ell_1$-norm, which is known as Basis Pursuit (BP) in the compressed sensing literature. Under a sparse true linear regression model with $p$ i.i.d. Gaussian features, we show that for a large range of $p$ up to a limit that grows exponentially with the number of samples $n$, with high probability the model error of BP is upper bounded by a value that decreases with $p$. To the best of our knowledge, this is the first analytical result in the literature establishing the double-descent of overfitting BP for finite $n$ and $p$. Further, our results reveal significant differences between the double-descent of BP and min $\ell_2$-norm solutions. Specifically, the double-descent upper-bound of BP is independent of the signal strength, and for high SNR and sparse models the descent-floor of BP can be much lower and wider than that of min $\ell_2$-norm solutions.

研究の動機と目的

  • $p > n$ の過学習状態における、$\ell_1$-ノルムを最小化する Basis Pursuit (BP) の一般化性能を分析すること。
  • 有限の $n$ および $p$ に対して、BP のモデル誤差に対する初めての非漸近的かつ高確率的な上界を確立すること。
  • BP の二重降下挙動を最小 $\ell_2$-ノルム解と比較し、信号強度依存性および降下フロアの幅の観点から検討すること。
  • $\ell_1$-ベースの過学習解(スパarsityを促進)が、完全な訓練適合を達成しても、深層ニューラルネットワークと同様に良好に一般化できるかどうかを調査すること。
  • BP が高SNRおよび高スパarsity条件下でも、低く広い降下フロアを持つ二重降下曲線を示す条件を解明すること。

提案手法

  • $p$ 個の i.i.d. ガウス特徴量と真のモデルのスパarsity $s$ を持つスパース線形回帰モデルを用いる。
  • Basis Pursuit (BP) を、正確なデータ適合(補間)を満たす $\ell_1$-ノルム最小化の解として分析する。
  • 濃度不等式とカイ二乗尾部バウンドを用いて、解の安定性に関連する $\min_i |\mathbf{A}_{i1}|$ の高確率的下界を導出する。
  • ガウス尾部推定と対数近似を用いて、確率的ベクトルの最大絶対値の逆数のバウンドを求める。
  • 高確率的下界 $\Pr\left(\frac{1}{\max_i |\mathbf{A}_{i1}|} \geq k\right)$ を導出し、これにより $\ell_1$-ノルムの解の制御を行う。
  • これらの確率的バウンドを設計行列の構造と組み合わせ、$p$ とともに減少する非漸近的モデル誤差の上界を導出する。

実験結果

リサーチクエスチョン

  • RQ1過パラメータ化状態 $p > n$ において、Basis Pursuit は二重降下の一般化誤差挙動を示すか?
  • RQ2信号強度依存性および降下フロアの幅の観点から、BP の二重降下曲線は最小 $\ell_2$-ノルム解とどのように異なるか?
  • RQ3$\ell_1$-ノルム最小化による過学習が、完全な訓練適合を達成しても低一般化誤差をもたらすことができるか?
  • RQ4BP が高確率で低モデル誤差を維持できる $p$ の最大範囲($n$ に対する相対的スケール)はどの程度か?
  • RQ5高SNR条件下でも、BP の二重降下挙動は信号強度に依存しないのか?

主な発見

  • $p$ が $\exp(\Theta(n))$ の広い範囲にわたり、Basis Pursuit のモデル誤差は高確率で $p$ とともに減少する値で上界が与えられる。
  • BP の二重降下上界は、信号強度に依存しないが、最小 $\ell_2$-ノルム解とは異なり、SNR に依存する。
  • 高SNRおよびスパースモデルの下では、BP の降下フロアは最小 $\ell_2$-ノルム解よりも著しく低く広い。
  • 本分析により、$p$ が増加するにつれて減少する $\ell_1$-ノルムの解に対する非漸近的かつ高確率的なバウンドが確立された。
  • 条件 $p - s \leq e^{(n-1)/16}/n$ の下で、解ベクトルの最大絶対値の逆数が閾値 $k$ を超える確率は $1 - 3/n$ 以上に下界される。
  • 導出されたバウンドにより、$p$ が $n$ に対して指数関数的に増加しても、BP が過学習状態においても低誤差を維持することが確認され、過パラメータ化に対するロバスト性が裏付けられた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。