Skip to main content
QUICK REVIEW

[論文レビュー] Robustness in sparse linear models: relative efficiency based on robust approximate message passing

Jelena Bradić|arXiv (Cornell University)|Jul 31, 2015
Statistical Methods and Inference参考文献 36被引用数 10
ひとこと要約

本稿は、$ p \gg n $ の高次元線形モデルに対して、重み付きでない、微分不能な損失関数を用いることで、重尾誤差下での推定効率を向上させる、頑健でスパースな近似メッセージパッシングアルゴリズム(RAMP)を導入する。重尾分布下でも、正則化付きの絶対偏差最小二乗(LAD)が最小二乗(LS)を効率で上回ることを確立し、非正規誤差下においてもスパarsityが存在する場合に、古典的なパターンが逆転することを示す。

ABSTRACT

Understanding efficiency in high dimensional linear models is a longstanding problem of interest. Classical work with smaller dimensional problems dating back to Huber and Bickel has illustrated the benefits of efficient loss functions. When the number of parameters $p$ is of the same order as the sample size $n$, $p \approx n$, an efficiency pattern different from the one of Huber was recently established. In this work, we consider the effects of model selection on the estimation efficiency of penalized methods. In particular, we explore whether sparsity, results in new efficiency patterns when $p > n$. In the interest of deriving the asymptotic mean squared error for regularized M-estimators, we use the powerful framework of approximate message passing. We propose a novel, robust and sparse approximate message passing algorithm (RAMP), that is adaptive to the error distribution. Our algorithm includes many non-quadratic and non-differentiable loss functions. We derive its asymptotic mean squared error and show its convergence, while allowing $p, n, s o \infty$, with $n/p \in (0,1)$ and $n/s \in (1,\infty)$. We identify new patterns of relative efficiency regarding a number of penalized $M$ estimators, when $p$ is much larger than $n$. We show that the classical information bound is no longer reachable, even for light--tailed error distributions. We show that the penalized least absolute deviation estimator dominates the penalized least square estimator, in cases of heavy--tailed distributions. We observe this pattern for all choices of the number of non-zero parameters $s$, both $s \leq n$ and $s \approx n$. In non-penalized problems where $s =p \approx n$, the opposite regime holds. Therefore, we discover that the presence of model selection significantly changes the efficiency patterns.

研究の動機と目的

  • 高次元スパース線形モデル($ p \approx n $ または $ p \gg n $)における推定効率を理解すること、特に正規分布でない誤差分布下での状況を対象とする。
  • 誤差が正規性から逸脱する場合、かつスパarsityが存在する場合に、古典的な正則化付きM推定量の頑健性の欠如を解決すること。
  • 一般損失関数(微分不能なものも含む)下での正則化付きM推定量の漸近的平均二乗誤差(AMSE)を導出すること。
  • モデル選択(スパarsity)が、古典的Huber型M推定量で観察される古典的効率パターンにどのように影響を与えるかを調査すること。
  • 高次元で$ p,n,s \to \infty $、$ n/p \in (0,1) $、$ n/s \in (1,\infty) $ の下で、多様な誤差分布に適応可能で、収束性を有する頑健で適応的なアルゴリズム(RAMP)を構築すること。

提案手法

  • 未知の誤差分布に適応可能な一般M推定量損失関数を用いる、新しい頑健な近似メッセージパッシング(RAMP)アルゴリズムを提案する。
  • 近似メッセージパッシング(AMP)フレームワークを用いて、高次元漸近的条件下での正則化付きM推定量の漸近的平均二乗誤差(AMSE)を導出する。
  • i.i.d. 正規設計行列とサブガウスノイズを仮定し、推定誤差のダイナミクスを追跡するための状態遷移解析を適用する。
  • 非二乗および非微分可能な損失関数(例:$ \rho(u) = |u| $ によるLAD)を組み込み、最小二乗法を超えて頑健性を向上させる。
  • 高次元で$ p,n,s \to \infty $、$ n/p \in (0,1) $、$ n/s \in (1,\infty) $ の下でRAMPアルゴリズムの収束を導出し、高次元領域における一貫性を保証する。
  • 信号対雑音比に基づくしきい値処理と、状態遷移におけるソフトしきい値処理を用いて、スパarsityを維持し、誤差伝搬を制御する。

実験結果

リサーチクエスチョン

  • RQ1高次元モデル($ p \gg n $)におけるスパarsityが、低次元設定と比較して、M推定量の古典的効率パターンにどのように影響を与えるか?
  • RQ2重尾誤差分布下で、$ p \gg n $ の場合に、頑健でスパースなM推定量が最小二乗法を上回る推定効率を達成できるか?
  • RQ3スパarsityが存在する場合($ s \approx n $ であっても)、正則化付き絶対偏差最小二乗(LAD)推定量が正則化付き最小二乗(LS)推定量を相対効率で上回るか?
  • RQ4高次元スパースモデル($ p \gg n $)において、軽尾誤差分布であっても、古典的情報量の下限は達成可能か?
  • RQ5モデル選択(スパarsity)の存在が、非正則化設定では通常LSがLADを上回るのに対し、逆転する効率支配パターンをどのように変化させるか?

主な発見

  • 重尾誤差分布下で、すべての$ s $の値($ s \leq n $ および $ s \approx n $ を含む)において、正則化付き絶対偏差最小二乗(LAD)推定量が正則化付き最小二乗(LS)推定量を相対効率で上回る。
  • スパarsityと高次元性の相互作用のため、高次元スパースモデルでは、軽尾誤差分布であっても古典的情報量の下限は達成不可能である。
  • スパarsityは効率パターンを根本的に変える:非正則化モデル($ s = p \approx n $)ではLSがLADを上回るが、正則化付きスパースモデル($ s \ll p $)ではLADがLSを上回る。
  • 提案されたRAMPアルゴリズムは、$ p,n,s \to \infty $ で$ n/p \in (0,1) $ および$ n/s \in (1,\infty) $ の下で収束し、漸近的平均二乗誤差(AMSE)の一貫性のある近似を提供する。
  • RAMPアルゴリズムは、未知の誤差分布に頑健かつ適応可能であり、最小二乗法を越える広範な非二乗および非微分可能な損失関数のクラスを扱える。
  • 状態遷移解析により、RAMPアルゴリズムが安定した誤差ダイナミクスを維持し、適切な条件下でソフトしきい値処理により最適なしきい値動作を達成し、真のスパース信号に収束することが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。