Skip to main content
QUICK REVIEW

[论文解读] Constructing confidence sets after lasso selection by randomized estimator augmentation

Seunghyun Min, Qing Zhou|arXiv (Cornell University)|Apr 16, 2019
Statistical Methods and Inference参考文献 15被引用 4
一句话总结

本文提出了一种新颖的方法,通过随机化估计量增强和从给定套索活跃集的响应条件分布进行马尔可夫链蒙特卡洛(MCMC)采样,构建套索选择后的有效联合置信集。通过在估计均值向量上引入随机化,该方法实现了具有保证覆盖概率的精确、低体积置信集,在覆盖性和效率方面优于现有最先进方法。

ABSTRACT

Although a few methods have been developed recently for building confidence intervals after model selection, how to construct confidence sets for joint post-selection inference is still an open question. In this paper, we develop a new method to construct confidence sets after lasso variable selection, with strong numerical support for its accuracy and effectiveness. A key component of our method is to sample from the conditional distribution of the response $y$ given the lasso active set, which, in general, is very challenging due to the tiny probability of the conditioning event. We overcome this technical difficulty by using estimator augmentation to simulate from this conditional distribution via Markov chain Monte Carlo given any estimate $ ildeμ$ of the mean $μ_0$ of $y$. We then incorporate a randomization step for the estimate $ ildeμ$ in our sampling procedure, which may be interpreted as simulating from a posterior predictive distribution by averaging over the uncertainty in $μ_0$. Our Monte Carlo samples offer great flexibility in the construction of confidence sets for multiple parameters. Extensive numerical results show that our method is able to construct confidence sets with the desired coverage rate and, moreover, that the diameter and volume of our confidence sets are substantially smaller in comparison with a state-of-the-art method.

研究动机与目标

  • 为解决在数据驱动的套索模型选择后构建多个参数联合置信集这一开放问题。
  • 克服在高维情况下,给定套索活跃集的响应条件分布采样计算困难的问题,该分布概率极低。
  • 开发一种灵活、可扩展的后模型选择推断框架,保持有效覆盖的同时最小化置信集体积。
  • 将推断范围从单个参数扩展到所选变量任意子集的联合推断。
  • 通过在均值估计不确定性上进行随机化平均,相比现有方法(如Lee等,2016年)提高数值稳定性和效率。

提出的方法

  • 使用估计量增强技术模拟给定套索活跃集 $ A $ 时响应 $ y $ 的条件分布,从而在复杂约束下实现MCMC采样。
  • 采用两阶段MCMC过程:首先,从给定 $ A $ 和固定估计值 $ ilde{ u} $ 的 $ y $ 条件分布中采样;然后,对 $ ilde{ u} $ 进行随机化,以近似后验预测分布。
  • 在估计均值 $ ilde{ u} $ 上引入随机化步骤,使能够对 $ u_0 $ 的不确定性进行平均,从而提高鲁棒性和覆盖性能。
  • 通过利用从联合条件分布中获得的蒙特卡洛样本构建置信集,实现对所选参数任意子集 $ u_B $ 的灵活推断。
  • 使用针对套索活跃集和变量选择阈值所诱导的多面体约束量身定制的提议的梅特罗波利斯-黑斯廷斯算法。
  • 通过在MCMC过程中使用截断正态提议和动态重新配置活跃集 $ ilde{eta}_j $,处理变量活跃集的变化。

实验结果

研究问题

  • RQ1当活跃集为数据驱动且高维时,能否在套索选择后构建有效的联合置信集?
  • RQ2如何高效地从给定套索活跃集的响应条件分布中采样,该分布在高维情况下概率极低?
  • RQ3与现有方法相比,随机化估计量增强能否提高后模型选择推断的数值稳定性和效率?
  • RQ4所提出的方法是否能产生比现有最先进方法(如Lee等,2016年)具有更好覆盖性和更小体积的置信集?
  • RQ5该方法能否推广以处理组结构或非套索选择程序?

主要发现

  • 所提方法在高维设置($ p > n $)下仍能实现名义覆盖率,验证了其有效性。
  • 数值结果表明,置信集的直径和体积显著小于Lee等(2016年)方法产生的结果,表明其精度更高。
  • 即使所选模型并非真实模型,该方法仍能保持有效覆盖,因其仅基于活跃集进行条件推断,不假设存在真实线性模型。
  • 对估计均值 $ ilde{ u} $ 的随机化显著提高了数值稳定性,并降低了置信集估计的方差。
  • MCMC采样器通过动态调整活跃集并使用截断提议,成功导航了条件分布复杂不规则的支持集。
  • 该方法可自然推广至任意子集 $ B \neq \text{all} $,避免了对同时个体区间采用过于保守的族错误率控制。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。