Skip to main content
QUICK REVIEW

[論文レビュー] Derivatives and residual distribution of regularized M-estimators with application to adaptive tuning

Pierre Bellec, Yiwei Shen|arXiv (Cornell University)|Jul 11, 2021
Statistical Methods and Inference参考文献 17被引用数 5
ひとこと要約

本稿は、ガウス型デザインと任意のノイズを伴う高次元線形モデルにおける正則化M推定量のための新しい適応的チューニング基準を開発する。応答変数および設計行列に関してM推定量のヤコビアンの正確な表現を導出し、残差分布を特徴づけ、ノイズ分布や設計共分散の知識を必要とせずに予測誤差の外側近似を可能にする基準を提案する。これにより、ロバストでデータ駆動型のチューニングが実現される。

ABSTRACT

This paper studies M-estimators with gradient-Lipschitz loss function regularized with convex penalty in linear models with Gaussian design matrix and arbitrary noise distribution. A practical example is the robust M-estimator constructed with the Huber loss and the Elastic-Net penalty and the noise distribution has heavy-tails. Our main contributions are three-fold. (i) We provide general formulae for the derivatives of regularized M-estimators $\hatβ(y,X)$ where differentiation is taken with respect to both $y$ and $X$; this reveals a simple differentiability structure shared by all convex regularized M-estimators. (ii) Using these derivatives, we characterize the distribution of the residual $r_i = y_i-x_i^ op\hatβ$ in the intermediate high-dimensional regime where dimension and sample size are of the same order. (iii) Motivated by the distribution of the residuals, we propose a novel adaptive criterion to select tuning parameters of regularized M-estimators. The criterion approximates the out-of-sample error up to an additive constant independent of the estimator, so that minimizing the criterion provides a proxy for minimizing the out-of-sample error. The proposed adaptive criterion does not require the knowledge of the noise distribution or of the covariance of the design. Simulated data confirms the theoretical findings, regarding both the distribution of the residuals and the success of the criterion as a proxy of the out-of-sample error. Finally our results reveal new relationships between the derivatives of $\hatβ(y,X)$ and the effective degrees of freedom of the M-estimator, which are of independent interest.

研究の動機と目的

  • ノイズ分布が未知で、かつ重尾である可能性がある状況において、理論的裏付けのある適応的基準を用いて正則化M推定量のチューニングパラメータを選択すること。
  • 勾配リプシッツ損失関数と凸罰則のもとで、正則化M推定量のヤコビアンを応答ベクトルyおよび設計行列Xに関して導出すること。
  • n ≈ p の中間的高次元的状態における残差の分布を特徴づけること。
  • ノイズ分布や設計共分散構造の知識を必要とせず、外側予測誤差を定数項を除いて近似するチューニング基準を構築すること。

提案手法

  • 勾配リプシッツ損失関数と凸罰則のもとで、M推定量のヤコビアンを応答変数yおよび設計行列Xに関して導出し、普遍的な微分可能性構造を明らかにする。
  • 導出したヤコビアンを用いて、有効自由度(df)および残差共分散構造を、行列V = diag(ψ′(r))(I - X(∂β̂/∂y))を用いて計算する。
  • 外側誤差を定数項を除いて近似するチューニング基準 Crit(ρ,g) = ||r + (df̂ / tr[V])ψ(r)||² を提案する。ここでrは残差ベクトル、ψは損失関数の微分である。
  • Huber損失とElastic-Netを含む一般的な損失-罰則ペアについて、df̂ / tr[V] の閉形式表現が得られることを確立し、計算の効率性を実現する。
  • ガウス型およびラデマッハ型設計のもとで、重尾ノイズを伴うシミュレーションを通じて手法の妥当性を検証し、残差分布の近似精度と基準の有効性を確認する。

実験結果

リサーチクエスチョン

  • RQ1高次元線形モデルにおける正則化M推定量の応答ベクトルおよび設計行列に関しての微分の一般形は何か?
  • RQ2一般の凸損失関数および罰則のもとで、n ≈ p の中間的状態における残差ベクトルの分布はどのように特徴づけられるか?
  • RQ3ノイズ分布や設計共分散の知識なしに、外側予測誤差を近似するデータ駆動型基準を構築できるか?
  • RQ4M推定量の微分とその有効自由度の関係は何か?
  • RQ5ノイズが重尾的であるか、設計が非ガウス的である状況において、提案された基準はチューニングパラメータの選択においてどの程度効果を示すか?

主な発見

  • 正則化M推定量の応答変数yに関するヤコビアンは、損失関数の微分と罰則のヘッセ行列にのみ依存する普遍的な構造を有する。
  • 残差ベクトルr = y - Xβ̂ は、中間的高次元的状態で特徴づけられる分布に従い、行列Vと比df̂ / tr[V] が重要な役割を果たす。
  • 提案された基準 Crit(ρ,g) = ||r + (df̂ / tr[V])ψ(r)||² は、推定量に依存しない定数項を除いて外側誤差を近似する。これにより、モデル選択の有効な代理指標となる。
  • この基準は適応的である:ノイズ分布や設計共分散Σの知識を必要とせず、重尾誤差下でも良好に機能する。
  • シミュレーションにより、残差の分布が理論的モデルによりよく近似されていることが確認され、基準は外側誤差を最小化するチューニングパラメータを的確に選択する。
  • 本稿は、M推定量の微分とその有効自由度との間の新たな関係を明らかにし、比df̂ / tr[V] が残差構造への重要な接続を提供することを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。