Skip to main content
QUICK REVIEW

[論文レビュー] A Bootstrap Lasso + Partial Ridge Method to Construct Confidence Intervals for Parameters in High-dimensional Sparse Linear Models

Hanzhong Liu, Xin Xu|arXiv (Cornell University)|Jun 7, 2017
Statistical Methods and Inference参考文献 42被引用数 15
ひとこと要約

本稿では、弱いかつまたは破綻した beta-min 条件のもとで既存手法に欠ける点を補うために、高次元スパース線形モデルにおける係数の信頼区間を構築するためのブートストラップlasso+部分リッジ(LPR)法を提案する。従来のde-sparsified lassoと比較して、特に小さな非ゼロ係数に対して、信頼区間のカバレッジ確率を50%向上させ、区間長を35%短縮する。

ABSTRACT

Constructing confidence intervals for the coefficients of high-dimensional sparse linear models remains a challenge, mainly because of the complicated limiting distributions of the widely used estimators, such as the lasso. Several methods have been developed for constructing such intervals. Bootstrap lasso+ols is notable for its technical simplicity, good interpretability, and performance that is comparable with that of other more complicated methods. However, bootstrap lasso+ols depends on the beta-min assumption, a theoretic criterion that is often violated in practice. Thus, we introduce a new method, called bootstrap lasso+partial ridge, to relax this assumption. Lasso+partial ridge is a two-stage estimator. First, the lasso is used to select features. Then, the partial ridge is used to refit the coefficients. Simulation results show that bootstrap lasso+partial ridge outperforms bootstrap lasso+ols when there exist small, but nonzero coefficients, a common situation that violates the beta-min assumption. For such coefficients, the confidence intervals constructed using bootstrap lasso+partial ridge have, on average, $50\%$ larger coverage probabilities than those of bootstrap lasso+ols. Bootstrap lasso+partial ridge also has, on average, $35\%$ shorter confidence interval lengths than those of the de-sparsified lasso methods, regardless of whether the linear models are misspecified. Additionally, we provide theoretical guarantees for bootstrap lasso+partial ridge under appropriate conditions, and implement it in the R package "HDCI."

研究の動機と目的

  • 従来の手法が複雑な漸近的分布を示すため失敗する高次元スパース線形モデルにおける係数の信頼区間を信頼性を持って構築する課題に対処する。
  • ブートストラップlasso+OLSが要求する制限の厳しい beta-min 条件を克服する。実際には小さな非ゼロ係数が存在する場合、この条件はしばしば満たされない。
  • モデルが誤って指定されている場合や小さな効果が存在する場合でも、高いカバレッジ確率と短い信頼区間長を維持する手法を開発する。
  • 実世界の応用(例:ゲノム研究)に適した、計算が単純で解釈可能かつ並列処理可能な推論手順を提供する。
  • モデルの誤指定に対して頑健であり、実データ(例:遺伝子発現研究)において小さな生物学的に意味のある効果を改善して検出できることを示す。

提案手法

  • 最初の段階でlassoを用いてモデル選択を行い、活性な予測変数を特定する。
  • 2番目の段階で部分リッジ回帰を適用し、選択されたモデル上で係数を再適合させ、分散を低減し推定の安定性を向上させる。
  • 2段階推定量を非パラメトリックブートストラップと組み合わせ、係数の標本分布を近似する。
  • lasso+部分リッジ推定量の百分位ブートストラップ分位点を用いて信頼区間を構築する。
  • リッジ回帰のバイアス低減特性を活用し、弱いスパarsityおよび小さな非ゼロ係数の下でも推論性能を向上させる。
  • 実用性と再現可能性を高めるために、Rパッケージ'HDCI'として実装する。

実験結果

リサーチクエスチョン

  • RQ1lassoと部分リッジを組み合わせた2段階推定量は、弱いかつbeta-min条件が満たされない状況下で、ブートストラップlasso+OLSと比較して信頼区間のカバレッジと長さを改善できるか?
  • RQ2小さな非ゼロ係数が存在する状況で、ブートストラップlasso+部分リッジ法のカバレッジ確率と区間長はどのように振る舞うか?
  • RQ3特に高次元設定下で、モデルが誤って指定されている場合でも、本手法は良好な性能を維持できるか?
  • RQ4真のモデルが誤って指定されている場合、本手法はde-sparsified lassoと比較して区間長とカバレッジにおいてどのように異なるか?
  • RQ5本手法は、遺伝子発現解析のような実世界の応用において、小さな生物学的に意味のある係数を効果的に同定できるか?

主な発見

  • ブートストラップlasso+部分リッジは、beta-min仮定に反する小さな非ゼロ係数に対して、平均でブートストラップlasso+OLSと比較して50%高いカバレッジ確率を達成する。
  • 本手法は、モデルの誤指定にかかわらず、de-sparsified lasso法と比較して平均で35%短い信頼区間を生成する。
  • 実際のfMRIデータ応用において、本手法は3つのde-sparsified lasso手法と比較して、より生物学的に解釈可能な遺伝子を同定した。これは、モデルの誤指定に対しても頑健であることを示唆する。
  • 本手法は、小さなが有意な係数を検出する点でブートストラップlasso+OLSを上回り、微細な調節効果を持つ機能的に重要な遺伝子を同定するのに適している。
  • すべてのスパースモデルに一貫して有効であるとは限らないが、本手法は実証的に優れた性能を示し、解釈可能性と計算の単純さという実用的利点を持つ。
  • 本手法は、主に大きな効果ではなく、ゲノム的・生物学的システムで一般的に見られる小さな非ゼロ係数を同定することに特に効果的である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。