Skip to main content
QUICK REVIEW

[Paper Review] The distribution of a linear predictor after model selection: Unconditional finite-sample distributions and asymptotic approximations

Hannes Leeb|RePEc: Research Papers in Economics|Nov 7, 2006
Statistical Methods and Inference4 citations
TL;DR

This paper derives the exact finite-sample distribution of a linear predictor after data-driven model selection in linear regression, providing a computable cumulative distribution function (cdf) and a simpler asymptotic approximation. The key contribution is characterizing the weak limit of this distribution under local alternatives, showing non-uniform convergence and identifying all possible accumulation points, which reveals significant distortions due to model selection beyond standard normality.

ABSTRACT

We analyze the (unconditional) distribution of a linear predictor that is constructed after a data-driven model selection step in a linear regression model. First, we derive the exact finite-sample cumulative distribution function (cdf) of the linear predictor, and a simple approximation to this (complicated) cdf. We then analyze the large-sample limit behavior of these cdfs, in the fixed-parameter case and under local alternatives.

Motivation & Objective

  • To derive the exact finite-sample cumulative distribution function (cdf) of a linear predictor after model selection in linear regression.
  • To develop a simple, uniform asymptotic approximation to the complex finite-sample cdf.
  • To analyze the large-sample limit behavior of the unconditional cdf under both fixed parameters and local alternatives.
  • To characterize all weak limit points of the post-model-selection cdf along sequences of parameters, especially under local alternatives.
  • To study model selection probabilities in finite and large samples, particularly under local parameter sequences.

Proposed method

  • Derives the exact finite-sample cdf, denoted $ G_{n,\theta,\sigma}(t) $, for a linear predictor after model selection in a normal linear model.
  • Introduces an 'idealized' cdf $ G^*_{n,\theta,\sigma}(t) $, assuming known error variance $ \sigma^2 $, to simplify analysis and enable asymptotic approximation.
  • Uses a general-to-specific model selection procedure based on sequential hypothesis tests to determine candidate models.
  • Applies weak convergence theory to characterize accumulation points of $ G_{n,\theta(n),\sigma(n)}(t) $ along sequences $ \theta(n), \sigma(n) $, especially under local alternatives.
  • Employs expansions of model selection probabilities and conditional cdfs to derive the asymptotic limit of the unconditional cdf.
  • Relies on results from Leeb [1] and Leeb & Pötscher [3,4] to establish convergence properties and non-uniformity in limit behavior.

Experimental results

Research questions

  • RQ1What is the exact finite-sample distribution of a linear predictor after model selection in a linear regression model with normal errors?
  • RQ2How does the unconditional distribution of the post-model-selection estimator differ from the standard normal distribution due to model selection?
  • RQ3What is the asymptotic behavior of the unconditional cdf under local alternatives, and how does it differ from the fixed-parameter case?
  • RQ4What are all possible weak limit points of the post-model-selection cdf along sequences of parameters $ \theta(n), \sigma(n) $?
  • RQ5How do model selection probabilities behave asymptotically, especially under local alternatives?

Key findings

  • The exact finite-sample cdf $ G_{n,\theta,\sigma}(t) $ is derived in closed form, though complex, and depends on the model selection path.
  • The asymptotic approximation $ G^*_{n,\theta,\sigma}(t) $, based on known $ \sigma^2 $, is uniformly close to $ G_{n,\theta,\sigma}(t) $ in large samples.
  • The unconditional cdf exhibits non-uniform convergence to its large-sample limit, necessitating analysis along parameter sequences.
  • All weak limit points of $ G_{n,\theta(n),\sigma(n)}(t) $ are characterized, and it is shown that local alternatives suffice to capture all possible limits.
  • The limit distribution is a convex combination of conditional cdfs weighted by asymptotic model selection probabilities, with non-degenerate limits even when $ \theta_{p^*} $ is infinite.
  • The limit cdf in (5.1) is a mixture of normal and non-normal components, reflecting the impact of model selection on inference.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.