Skip to main content
QUICK REVIEW

[论文解读] Can One Estimate The Unconditional Distribution of Post-Model-Selection Estimators?

Hannes Leeb, Benedikt M. Pötscher|ArXiv.org|Apr 12, 2007
Statistical Methods and Inference被引用 4
一句话总结

本文证明,即使在渐近情况下,要以合理精度估计模型选择后估计量的无条件分布也是根本不可能的。它证明了由于估计误差的极小化最大下界趋近于1/2或1,任何估计量都无法实现一致(甚至局部一致)性,这使得在标准准则下进行模型选择后的可靠推断变得不可行。

ABSTRACT

We consider the problem of estimating the unconditional distribution of a post-model-selection estimator. The notion of a post-model-selection estimator here refers to the combined procedure resulting from first selecting a model (e.g., by a model selection criterion like AIC or by a hypothesis testing procedure) and then estimating the parameters in the selected model (e.g., by least-squares or maximum likelihood), all based on the same data set. We show that it is impossible to estimate the unconditional distribution with reasonable accuracy even asymptotically. In particular, we show that no estimator for this distribution can be uniformly consistent (not even locally). This follows as a corollary to (local) minimax lower bounds on the performance of estimators for the distribution; performance is here measured by the probability that the estimation error exceeds a given threshold. These lower bounds are shown to approach 1/2 or even 1 in large samples, depending on the situation considered. Similar impossibility results are also obtained for the distribution of linear functions (e.g., predictors) of the post-model-selection estimator.

研究动机与目标

  • 研究模型选择后估计量的无条件分布是否可以被一致估计。
  • 解决模型选择后推断的关键问题,即传统统计理论因数据驱动的模型不确定性而失效。
  • 确定尽管该分布对未知参数具有复杂依赖性,是否仍可实现分布估计的统一或局部一致性。
  • 将不可能性结果扩展至模型选择后估计量的线性函数,如预测器。
  • 挑战基于渐近近似构建的插补估计量的有效性,这些估计量可能因非一致收敛而存在问题。

提出的方法

  • 推导任意估计量在模型选择后估计量无条件分布上的局部极小化最大下界。
  • 使用总变差距离度量估计误差,并建立误差超过阈值的概率的界限。
  • 应用Leeb和Pötscher(2006a)关于渐近一致连续性和连续性的结果,分析收敛行为。
  • 考虑基于AIC或假设检验等准则的模型选择程序,随后通过最小二乘法或最大似然法进行估计。
  • 在不同参数配置下分析模型选择后估计量的分布,特别是边界情况附近。
  • 将分析扩展至估计量的线性函数(如预测器),并证明类似的不可能性结果。

实验结果

研究问题

  • RQ1即使在渐近情况下,模型选择后估计量的无条件分布是否可以被一致估计?
  • RQ2模型选择后估计量分布的估计量是否可实现统一或局部一致性?
  • RQ3基于渐近近似的插补估计量是否能克服有限样本分布中固有的非一致性问题?
  • RQ4此类分布估计量的极小化最大误差界是多少?
  • RQ5对于模型选择后估计量的线性函数(如预测器),类似的不可能性结果是否仍然成立?

主要发现

  • 任何估计模型选择后估计量无条件分布的估计量都无法实现统一一致,甚至无法实现局部一致。
  • 估计误差超过阈值的概率的极小化最大下界在大样本中趋近于1/2,甚至趋近于1。
  • 该不可能性结果与所用模型选择准则无关(例如AIC、假设检验),表明存在根本性限制。
  • 该不可能性同样适用于模型选择后估计量的线性函数(如预测器)的分布。
  • 有限样本分布向其渐近极限的非一致收敛性,削弱了插补估计量的可靠性。
  • 结果表明,通过一致估计无条件抽样分布,无法实现模型选择后的有效推断。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。