Skip to main content
QUICK REVIEW

[论文解读] Perturbation Bootstrap in Adaptive Lasso

Debraj Das, Karl Gregory|arXiv (Cornell University)|Mar 9, 2017
Statistical Methods and Inference参考文献 24被引用 3
一句话总结

本文提出了一种改进的扰动自展法(perturbation bootstrap)用于自适应Lasso(Alasso),在模型维度随样本量增长的高维设定下实现了二阶正确性。与Minnier等人(2011)提出的朴素扰动自展法不同,后者无法校正高阶偏差,本文通过修改自展目标函数并引入学生化(studentization)方法,确保了Alasso估计量的分布近似更加准确,显著提升了相对于“理想正态近似”(oracle normal approximation)的推断精度。

ABSTRACT

The Adaptive Lasso(Alasso) was proposed by Zou [ extit{J. Amer. Statist. Assoc. extbf{101} (2006) 1418-1429}] as a modification of the Lasso for the purpose of simultaneous variable selection and estimation of the parameters in a linear regression model. Zou (2006) established that the Alasso estimator is variable-selection consistent as well as asymptotically Normal in the indices corresponding to the nonzero regression coefficients in certain fixed-dimensional settings. In an influential paper, Minnier, Tian and Cai [ extit{J. Amer. Statist. Assoc. extbf{106} (2011) 1371-1382}] proposed a perturbation bootstrap method and established its distributional consistency for the Alasso estimator in the fixed-dimensional setting. In this paper, however, we show that this (naive) perturbation bootstrap fails to achieve second order correctness in approximating the distribution of the Alasso estimator. We propose a modification to the perturbation bootstrap objective function and show that a suitably studentized version of our modified perturbation bootstrap Alasso estimator achieves second-order correctness even when the dimension of the model is allowed to grow to infinity with the sample size. As a consequence, inferences based on the modified perturbation bootstrap will be more accurate than the inferences based on the oracle Normal approximation. We give simulation studies demonstrating good finite-sample properties of our modified perturbation bootstrap method as well as an illustration of our method on a real data set.

研究动机与目标

  • 解决Minnier等人(2011)提出的朴素扰动自展法在高维设定下无法实现自适应Lasso估计量二阶正确性的局限性。
  • 开发一种计算高效的自展方法,准确近似当预测变量数量随样本量增长时Alasso估计量的有限样本分布。
  • 为稀疏高维线性模型中非零回归系数的推断,提供一种对“理想正态近似”的改进替代方案。
  • 在允许维度发散的弱正则性条件下,建立理论保证,特别是二阶正确性。

提出的方法

  • 提出一种改进的扰动自展目标函数,通过在自展损失函数中引入校正后的惩罚项,以校正Alasso估计量中的偏差。
  • 提出一种改进自展估计量的学生化版本,以确保其渐近分布与Alasso估计量的真实抽样分布等价。
  • 采用埃奇沃斯展开(Edgeworth expansion)技术和高阶渐近理论,分析自展分布向真实抽样分布的收敛速度。
  • 利用泰勒展开和鞅型论证,控制在高维渐近设定下自展方差与惩罚项的估计误差。
  • 应用霍夫丁(Hoeffding)和伯恩斯坦(Bernstein)不等式,在弱矩条件下方建立自展统计量的统一集中界。
  • 推导出改进自展分布与真实抽样分布的埃奇沃斯展开在 $ o(n^{-1}) $ 阶次内渐近等价,从而确保二阶正确性。

实验结果

研究问题

  • RQ1Minnier等人(2011)提出的朴素扰动自展法是否能在高维设定下实现对自适应Lasso估计量分布的二阶正确近似?
  • RQ2能否构造一种改进的扰动自展目标函数,使得当模型维度随样本量增长时,仍能实现二阶正确性?
  • RQ3在有限样本中,改进自展法与理想正态近似相比性能如何?
  • RQ4在何种理论条件下,改进自展法能在高维线性模型中实现二阶精度?

主要发现

  • 即使在固定维数设定下,朴素扰动自展法也无法实现对自适应Lasso估计量分布的二阶正确近似,原因在于惩罚项带来的偏差未被校正。
  • 所提出的改进扰动自展目标函数在分布近似中实现了二阶正确性,即使当预测变量数量随样本量增长时亦成立。
  • 当改进自展估计量经过适当的学生化处理后,其在埃奇沃斯展开中与Alasso估计量的有限样本分布一致至 $ o(n^{-1}) $ 项,从而显著提升了推断精度。
  • 模拟研究显示,改进自展法在有限样本中表现良好,其置信区间覆盖概率更接近名义水平,优于基于理想正态近似的推断方法。
  • 该方法计算高效,适用于高维设定,是高维线性模型中非零回归系数推断的实用替代方案,优于渐近正态近似。
  • 理论结果证实,改进自展分布以与真实估计量相同的速率收敛至真实抽样分布,且其埃奇沃斯展开在 $ o(n^{-1}) $ 阶次内一致,验证了其在推断中的有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。