Skip to main content
QUICK REVIEW

[论文解读] Valid Post-selection Inference in Assumption-lean Linear Regression

Arun Kumar Kuchibhotla, Lawrence D. Brown|arXiv (Cornell University)|Jun 11, 2018
Statistical Methods and Inference参考文献 24被引用 7
一句话总结

本文提出了一种计算高效的后选择推断方法,适用于线性回归中任意数据驱动的变量选择,利用确定性不等式构建的置信区域显著小于现有方法。该方法在一般条件下(包括依赖和非同分布数据)保持有效性,并在固定协变量下实现最优尺寸。

ABSTRACT

Construction of valid statistical inference for estimators based on data-driven selection has received a lot of attention in the recent times. Berk et al. (2013) is possibly the first work to provide valid inference for Gaussian homoscedastic linear regression with fixed covariates under arbitrary covariate/variable selection. The setting is unrealistic and is extended by Bachoc et al. (2016) by relaxing the distributional assumptions. A major drawback of the aforementioned works is that the construction of valid confidence regions is computationally intensive. In this paper, we first prove that post-selection inference is equivalent to simultaneous inference and then construct valid post-selection confidence regions which are computationally simple. Our construction is based on deterministic inequalities and apply to independent as well as dependent random variables without the requirement of correct distributional assumptions. Finally, we compare the volume of our confidence regions with the existing ones and show that under non-stochastic covariates, our regions are much smaller.

研究动机与目标

  • 通过在数据驱动模型选择后实现有效的统计推断,解决科学中的可重复性危机。
  • 克服经典推断的局限性,后者在建模决策依赖于数据时(研究者自由度)会失效。
  • 在模型误设下提供有效推断,且无需强分布假设。
  • 构建计算高效且体积小于现有方法的置信区域,尤其在固定协变量下表现更优。
  • 建立一种对数据依赖性和非同分布性具有鲁棒性的后选择推断框架。

提出的方法

  • 证明后选择推断等价于同时推断,从而实现对置信区域的统一处理。
  • 使用确定性不等式而非分布近似来构建置信区域,确保有效性,且无需正确模型假设。
  • 通过线性不等式组定义置信区域,利用线性规划实现局部推断的高效计算。
  • 利用普通最小二乘法(OLS)回归的估计方程框架,推导对模型误设具有鲁棒性的推断程序。
  • 将该方法应用于独立和依赖的随机变量,包括非随机(固定)协变量。
  • 证明置信区域在设计矩阵的任意线性变换下保持不变,尽管这可能与计算效率冲突。

实验结果

研究问题

  • RQ1是否可以不依赖于限制性分布假设或计算密集型方法,构建有效的后选择置信区域?
  • RQ2后选择置信区域的体积与现有方法相比如何,尤其在固定协变量下?
  • RQ3所提方法是否能在模型误设和数据依赖下保持有效性?
  • RQ4后选择推断中是否存在计算效率与仿射不变性之间的权衡?
  • RQ5该框架能否扩展至OLS线性回归以外的其他M-估计问题?

主要发现

  • 所提置信区域的体积显著小于Berk等人(2013)和Bachoc等人(2016)的方法,尤其在固定协变量下表现更优。
  • 置信区域在渐近下保持有效性,覆盖概率至少为1−α,即使在模型误设和依赖数据条件下亦成立。
  • 如注记4.6所示,在固定协变量下,该方法实现了最优尺寸,且未牺牲有效性。
  • 置信区域基于确定性不等式,确保有效性,无需独立同分布或正态性假设。
  • 该方法可通过线性规划实现对单个系数的局部推断,尽管此类区域可能偏保守。
  • 该框架揭示了后选择推断与高维稀疏诱导估计器之间的联系,提示其可推广至其他M-估计问题。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。