Skip to main content
QUICK REVIEW

[论文解读] Selective inference after variable selection via multiscale bootstrap

Yoshikazu Terada, Hidetoshi Shimodaira|arXiv (Cornell University)|May 25, 2019
Statistical Methods and Inference参考文献 3被引用 6
一句话总结

本文提出一种多尺度自展法,用于回归中变量选择后的选择性推断,解决了p值和置信区间的选择偏差问题。通过定义更具灵活性的选择事件——聚焦于变量是否被纳入模型而非特定模型——该方法在MCP等各类算法中均能实现有效的推断,且计算成本与经典自展法相当。

ABSTRACT

A general resampling approach is considered for selective inference problem after variable selection in regression analysis. Even after variable selection, it is important to know whether the selected variables are actually useful by showing $p$-values and confidence intervals of regression coefficients. In the classical approach, significance levels for the selected variables are usually computed by $t$-test but they are subject to selection bias. In order to adjust the bias in this post-selection inference, most existing studies of selective inference consider the specific variable selection algorithm such as Lasso for which the selection event can be explicitly represented as a simple region in the space of the response variable. Thus, the existing approach cannot handle more complicated algorithm such as MCP (minimax concave penalty). Moreover, most existing approaches set an event, that a specific model is selected, as the selection event. This selection event is too restrictive and may reduce the statistical power, because the hypothesis selection with a specific variable only depends on whether the variable is selected or not. In this study, we consider more appropriate selection event such that the variable is selected, and propose a new bootstrap method to compute an approximately unbiased selective $p$-value for the selected variable. Our method is applicable to a wide class of variable selection algorithms. In addition, the computational cost of our method is the same order as the classical bootstrap method. Through the numerical experiments, we show the usefulness of our selective inference approach.

研究动机与目标

  • 解决回归中变量选择后p值和置信区间的选择偏差问题。
  • 克服现有选择性推断方法的局限性,这些方法依赖于限制性较强的选择事件或特定算法(如Lasso)。
  • 开发一种可推广的方法,适用于MCP(最小凹惩罚)等复杂变量选择算法。
  • 通过采用聚焦于变量纳入而非精确模型选择的较宽松选择事件,提高统计检验效能。
  • 在保持推断有效性的同时,确保计算效率与经典自展法相当。

提出的方法

  • 将选择事件定义为特定变量被纳入模型,而非选择某一特定模型。
  • 提出一种多尺度自展框架,以近似给定选择事件下检验统计量的条件分布。
  • 通过重采样方法,在选择性推断框架下估计检验统计量的零分布。
  • 通过基于自展法的估计,对观测到的选择事件进行条件处理,构建近似无偏的p值。
  • 通过保持与经典自展法相同的复杂度量级,确保方法的计算效率。
  • 将该方法应用于广泛的一类变量选择算法,包括MCP等非凸惩罚方法。

实验结果

研究问题

  • RQ1当选择事件并非特定模型而是变量纳入时,如何在变量选择后可靠地进行选择性推断?
  • RQ2基于自展法的方法能否在多种变量选择算法中提供有效的p值和置信区间,以纠正选择偏差?
  • RQ3与经典自展法相比,该方法的计算成本如何?是否可保持相近?
  • RQ4与现有方法相比,该方法在统计效能和第一类错误控制方面表现如何?
  • RQ5该方法能否扩展至MCP等非凸惩罚方法?这些方法的选择事件复杂且难以明确刻画。

主要发现

  • 所提出的多尺度自展法可产生近似无偏的p值,有效纠正变量选择后推断中的选择偏差。
  • 该方法适用于广泛的一类变量选择算法,包括现有选择性推断方法难以处理的MCP。
  • 所提方法的计算成本与经典自展法处于同一数量级,具有良好的可扩展性和实用性。
  • 数值实验表明,该方法在保持适当的第一类错误率方面表现有效,并相比朴素的后选择推断显著提升了统计效能。
  • 基于变量纳入的选择事件可带来比基于特定模型的选择事件更高的推断效能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。