Skip to main content
QUICK REVIEW

[論文レビュー] On the Optimal Weighted $\ell_2$ Regularization in Overparameterized Linear Regression

Denny Wu, Ji Xu|arXiv (Cornell University)|Jun 10, 2020
Sparse and Compressive Sensing Techniques参考文献 45被引用数 17
ひとこと要約

この論文は、重み付き $\lambda_2$ 正則化を施した過パラメータ化線形回帰における予測リスクの正確な漸近的特徴付けを提供し、特異なデータおよび信号の共分散のもとで最適リッジパラメータ $\lambda_{\text{opt}}$ が負になる可能性があることを示している。リッジレスおよび最適正則化された設定の両方における最適重み行列 $\boldsymbol{\Sigma}_w$ を導出し、暗黙の正則化を裏付け、主成分回帰における二重降下現象を説明している。

ABSTRACT

We consider the linear model $\mathbf{y} = \mathbf{X} \mathbfβ_\star + \mathbfε$ with $\mathbf{X}\in \mathbb{R}^{n imes p}$ in the overparameterized regime $p>n$. We estimate $\mathbfβ_\star$ via generalized (weighted) ridge regression: $\hat{\mathbfβ}_λ= \left(\mathbf{X}^T\mathbf{X} + λ\mathbfΣ_w ight)^\dagger \mathbf{X}^T\mathbf{y}$, where $\mathbfΣ_w$ is the weighting matrix. Under a random design setting with general data covariance $\mathbfΣ_x$ and anisotropic prior on the true coefficients $\mathbb{E}\mathbfβ_\star\mathbfβ_\star^T = \mathbfΣ_β$, we provide an exact characterization of the prediction risk $\mathbb{E}(y-\mathbf{x}^T\hat{\mathbfβ}_λ)^2$ in the proportional asymptotic limit $p/n ightarrow γ\in (1,\infty)$. Our general setup leads to a number of interesting findings. We outline precise conditions that decide the sign of the optimal setting $λ_{ m opt}$ for the ridge parameter $λ$ and confirm the implicit $\ell_2$ regularization effect of overparameterization, which theoretically justifies the surprising empirical observation that $λ_{ m opt}$ can be negative in the overparameterized regime. We also characterize the double descent phenomenon for principal component regression (PCR) when both $\mathbf{X}$ and $\mathbfβ_\star$ are anisotropic. Finally, we determine the optimal weighting matrix $\mathbfΣ_w$ for both the ridgeless ($λ o 0$) and optimally regularized ($λ= λ_{ m opt}$) case, and demonstrate the advantage of the weighted objective over standard ridge regression and PCR.

研究の動機と目的

  • 一般の非等方的データおよび信号の共分散のもとで、過パラメータ化された状況($p > n$)における一般化リッジ回帰の予測リスクを特徴付けること。
  • 過パラメータ化された設定において経験的に観察された負の最適リッジパラメータのパラドックスを解明すること。
  • リッジレス($\lambda \to 0$)および最適に正則化された($\lambda = \lambda_{\text{opt}}$)状況の両方における最適重み行列 $\boldsymbol{\Sigma}_w$ を導出すること。
  • 非等方的条件下における主成分回帰の二重降下行動を説明すること。

提案手法

  • 比例的漸近的極限 $p/n \to \gamma \in (1, \infty)$ における予測リスクを、確率的行列理論および母集団固有値分布解析を用いて導出する。
  • 一般の正定値重み行列 $\boldsymbol{\Sigma}_w$ を持つ一般化リッジ推定量 $\hat{\boldsymbol{\beta}}_\lambda = (\boldsymbol{X}^\top\boldsymbol{X} + \lambda\boldsymbol{\Sigma}_w)^\dagger \boldsymbol{X}^\top\boldsymbol{y}$ を導入する。
  • i.i.d. 特徴量 $\boldsymbol{x}_i \sim \mathcal{N}(0, \boldsymbol{\Sigma}_x)$ と非等方的事前分布 $\mathbb{E}[\boldsymbol{\beta}_\star\boldsymbol{\beta}_\star^\top] = \boldsymbol{\Sigma}_\beta$ を持つランダム設計モデルを用いる。
  • 予測リスク $\mathbb{E}(y - \boldsymbol{x}^\top\hat{\boldsymbol{\beta}}_\lambda)^2$ の正確な表現を、$\boldsymbol{\Sigma}_x$ および $\boldsymbol{\Sigma}_\beta$ の固有値および正則化パラメータ $\lambda$ の関数として導出する。
  • リスク関数の微分の符号を分析することで、$\lambda_{\text{opt}} < 0$ となる条件を確立する。
  • 漸近的予測リスクを最小化することで最適 $\boldsymbol{\Sigma}_w$ を計算し、それが $\boldsymbol{\Sigma}_x$ および $\boldsymbol{\Sigma}_\beta$ の共同固有スペクトル構造に依存することを示す。

実験結果

リサーチクエスチョン

  • RQ1過パラメータ化された状況において、最適リッジパラメータ $\lambda_{\text{opt}}$ が負となる条件は何か?
  • RQ2データ共分散 $\boldsymbol{\Sigma}_x$ および信号共分散 $\boldsymbol{\Sigma}_\beta$ の非等方的性が、予測リスクおよび最適正則化に与える影響は何か?
  • RQ3リッジレスおよび最適に正則化された設定の両方において、予測リスクを最小化する最適重み行列 $\boldsymbol{\Sigma}_w$ は何か?
  • RQ4非等方的条件下で主成分回帰における二重降下現象はどのように現れるか?
  • RQ5重み付きリッジ回帰は、高次元の過パラメータ化された状況において、標準的リッジ回帰やPCRを上回る性能を示せるか?

主な発見

  • 信号およびデータの共分散 $\boldsymbol{\Sigma}_\beta$ と $\boldsymbol{\Sigma}_x$ が非等方的である場合、最適リッジパラメータ $\lambda_{\text{opt}}$ は負になり得る。これは経験的観察の理論的裏付けを提供する。
  • データ共分散 $\boldsymbol{\Sigma}_x$ と信号共分散 $\boldsymbol{\Sigma}_\beta$ が不整合な場合、主成分回帰における予測リスクは二重降下を示し、過パラメータ化領域に複数のピークを示す。
  • リッジレス回帰における最適重み行列 $\boldsymbol{\Sigma}_w$ は $\boldsymbol{\Sigma}_w = \left( (\boldsymbol{\Sigma}_x - c\mathbf{I})^2 + d\mathbf{I} \right)^{-1}$ で与えられ、ここで $c$ および $d$ は固有値の条件付き期待値に依存する。
  • 最適に正則化された回帰では、最適 $\boldsymbol{\Sigma}_w$ が $\boldsymbol{\Sigma}_x$ および $\boldsymbol{\Sigma}_\beta$ の共同固有スペクトル分布の関数として導出され、標準的リッジ回帰やPCRよりも一般化性能が向上する。
  • 本研究で導出した漸近的予測リスクは、離散的および連続的分布を含むさまざまな設定において、有限標本のシミュレーションと一致する。
  • 最適 $\boldsymbol{\Sigma}_w$ を用いた重み付きリッジ回帰は、標準的リッジ回帰やPCRよりも低い予測リスクを達成し、特に非等方的かつ不整合な設定において顕著に優れた性能を示す。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。