Skip to main content
QUICK REVIEW

[論文レビュー] Tractable Post-Selection Maximum Likelihood Inference for the Lasso

Amit Meir, Mathias Drton|arXiv (Cornell University)|May 26, 2017
Statistical Methods and Inference参考文献 41被引用数 5
ひとこと要約

本稿は、lassoで選択された高次元線形モデルにおける tractable な選択後最尤推定のための確率的最適化フレームワークを提案する。選択後スコア関数の不偏でノイズのある推定値を生成することにより、一貫性のある推定と妥当な信頼区間を可能にし、名目水準に近いカバレッジを達成する。これは、スパースな設定において、標準的なlassoやリファitted最小二乗法を上回る性能を示す。

ABSTRACT

Applying standard statistical methods after model selection may yield inefficient estimators and hypothesis tests that fail to achieve nominal type-I error rates. The main issue is the fact that the post-selection distribution of the data differs from the original distribution. In particular, the observed data is constrained to lie in a subset of the original sample space that is determined by the selected model. This often makes the post-selection likelihood of the observed data intractable and maximum likelihood inference difficult. In this work, we get around the intractable likelihood by generating noisy unbiased estimates of the post-selection score function and using them in a stochastic ascent algorithm that yields correct post-selection maximum likelihood estimates. We apply the proposed technique to the problem of estimating linear models selected by the lasso. In an asymptotic analysis the resulting estimates are shown to be consistent for the selected parameters and to have a limiting truncated normal distribution. Confidence intervals constructed based on the asymptotic distribution obtain close to nominal coverage rates in all simulation settings considered, and the point estimates are shown to be superior to the lasso estimates when the true model is sparse.

研究の動機と目的

  • モデル選択後の無効な推定を回避する根本的課題に取り組むこと、特にlassoを用いた高次元設定において。
  • lasso回帰における選択事象を考慮した最尤推定を計算可能な方法で開発すること。
  • 選択バイアスを補正することで、正確なカバレッジ率を達成する信頼区間を構築すること。
  • ガウス線形回帰にとどまらず、一般化線形モデルを含む指数型分布族モデルへも一般化可能なフレームワークを提供すること。

提案手法

  • 選択後スコア関数のノイズあり不偏推定値を用いた確率的勾配上昇法により、lasso選択後の最尤推定値を計算する。
  • 選択事象によって定義されるランダムな多面体集合にデータが制約されるため、選択後尤度は計算不能であるという事実を活用する。
  • 制御されたノイズを伴うモンテカルロサンプリングによりスコア関数を推定し、真の選択後MLEへの収束を可能にする。
  • 正則性条件の下で、推定値は漸近的に切断正規分布に収束するため、漸近的に有効である。
  • 推定値の漸近的分布を用いて信頼区間を構築し、正しいカバレッジを保証する。
  • 一貫性があり不偏なスコア推定器を用いた確率的勾配上昇アルゴリズムにより、フレームワークを実装する。

実験結果

リサーチクエスチョン

  • RQ1選択後尤度が計算不能であるにもかかわらず、lassoによるモデル選択後に有効な最尤推定値を計算できるか?
  • RQ2提案された確率的最適化手法は、選択後サンプリングの下で一貫性があり、漸近的に正規分布に従う推定値を生成するか?
  • RQ3得られた信頼区間は、標準的で補正のない区間と比較して、カバレッジとサイズの点でどのように異なるか?
  • RQ4本手法は、ガウス線形回帰にとどまらず、他の指数型分布族モデルへも拡張可能か?

主な発見

  • 提案された条件付き最尤推定値は、選択モデルのパラメータに対して一貫性があり、漸近的に切断正規分布に従う。
  • 漸近的分布に基づく信頼区間は、すべてのシミュレーション設定で95%に近いカバレッジを達成するが、補正のない区間は顕著に反保守的である。
  • 真のモデルがスパースな場合、点推定値は標準的lassoおよびリファitted最小二乗法を上回る予測精度を示す。
  • 条件付き・ワルド信頼区間は、ポリヘドラル区間と比較して著しく短く、ばらつきも小さいが、同程度のカバレッジを維持する。
  • 多くの共変数がある場合でも計算が可能であり、高次元設定における実用的な選択後推定を可能にする。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。