Skip to main content
QUICK REVIEW

[論文レビュー] Selective inference after variable selection via multiscale bootstrap

Yoshikazu Terada, Hidetoshi Shimodaira|arXiv (Cornell University)|May 25, 2019
Statistical Methods and Inference参考文献 3被引用数 6
ひとこと要約

本稿では、回帰における変数選択後の選択的仮説検定のためのマルチスケールブートストラップ法を提案する。この手法は、p値と信頼区間における選択バイアスを是正することを目的としており、特定のモデルではなく変数の包含に注目するより柔軟な選択イベントを定義することで、MCPなどの多様なアルゴリズムに対しても有効な推論を可能にする。計算コストは古典的ブートストラップと同等であり、実用的である。

ABSTRACT

A general resampling approach is considered for selective inference problem after variable selection in regression analysis. Even after variable selection, it is important to know whether the selected variables are actually useful by showing $p$-values and confidence intervals of regression coefficients. In the classical approach, significance levels for the selected variables are usually computed by $t$-test but they are subject to selection bias. In order to adjust the bias in this post-selection inference, most existing studies of selective inference consider the specific variable selection algorithm such as Lasso for which the selection event can be explicitly represented as a simple region in the space of the response variable. Thus, the existing approach cannot handle more complicated algorithm such as MCP (minimax concave penalty). Moreover, most existing approaches set an event, that a specific model is selected, as the selection event. This selection event is too restrictive and may reduce the statistical power, because the hypothesis selection with a specific variable only depends on whether the variable is selected or not. In this study, we consider more appropriate selection event such that the variable is selected, and propose a new bootstrap method to compute an approximately unbiased selective $p$-value for the selected variable. Our method is applicable to a wide class of variable selection algorithms. In addition, the computational cost of our method is the same order as the classical bootstrap method. Through the numerical experiments, we show the usefulness of our selective inference approach.

研究の動機と目的

  • 回帰における変数選択後のp値と信頼区間における選択バイアスを是正すること。
  • Lassoなどの特定のアルゴリズムに依存する、制限の厳しい選択イベントに依存する既存の選択的仮説検定手法の限界を克服すること。
  • MCP(ミニマックス凹型ペナルティ)のような複雑な変数選択アルゴリズムに一般化可能な手法を開発すること。
  • 特定のモデル選択ではなく変数の包含に注目するより制限の少ない選択イベントを用いることで、統計的検出力の向上を図ること。
  • 古典的ブートストラップと同等の計算効率を維持しつつ、推論の有効性を保つこと。

提案手法

  • 選択イベントを特定のモデルの選択ではなく、特定の変数がモデルに含まれることとして定義する。
  • テスト統計量の条件付き分布を近似するためのマルチスケールブートストラップフレームワークを提案する。
  • 再標本抽出を用いて、選択的仮説検定枠組み下での帰無分布を推定する。
  • ブートストラップに基づく推定を用いて、観察された選択イベントを条件として、ほぼ不偏なp値を構築する。
  • 古典的ブートストラップと同等の計算量のオーダーを維持することで、計算効率を確保する。
  • 凸でない罰則(MCPを含む)を含む広範な変数選択アルゴリズムにこの手法を適用する。

実験結果

リサーチクエスチョン

  • RQ1選択イベントが特定のモデルではなく変数の包含である場合、変数選択後の選択的仮説検定を信頼性を持って行うにはどうすればよいか?
  • RQ2ブートストラップに基づく手法が、多様な変数選択アルゴリズムにわたってバイアスを是正する有効なp値と信頼区間を提供できるか?
  • RQ3このような手法の計算コストは古典的ブートストラップと比較してどの程度か。同等に保てると考えられるか?
  • RQ4既存の手法と比較して、提案手法の統計的検出力と第一種過誤の制御性能はいかがであるか?
  • RQ5選択イベントが複雑で明確に特徴づけにくいMCPのような非凸罰則に対しても、この手法を拡張可能か?

主な発見

  • 提案されたマルチスケールブートストラップ法は、変数選択後の推論における選択バイアスを是正するほぼ不偏なp値を生成する。
  • この手法は、Lassoとは異なり、容易に扱えないMCPを含む広範な変数選択アルゴリズムに適用可能である。
  • 提案手法の計算コストは古典的ブートストラップと同等のオーダーであり、スケーラブルで実用的である。
  • 数値実験により、この手法が適切な第一種過誤率を維持するとともに、単純な選択後推論に比べて統計的検出力を向上させることを示している。
  • 変数包含に基づく選択イベントの使用は、モデル固有の選択イベントよりもより高い検出力をもたらす。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。