Skip to main content
QUICK REVIEW

[Paper Review] Goodness of fit tests for high-dimensional models

Rajen D. Shah, Peter Bühlmann|arXiv (Cornell University)|Nov 10, 2015
Statistical Methods and Inference28 references3 citations
TL;DR

This paper introduces Residual Prediction (RP) tests, a unified framework for goodness-of-fit testing in low- and high-dimensional linear models. By regressing scaled residuals from OLS or Lasso fits using prediction error as a test statistic, and leveraging simulation or parametric bootstrap for critical values, RP tests effectively detect model misspecifications such as heteroscedasticity and nonlinearity while outperforming state-of-the-art methods for variable significance testing.

ABSTRACT

In this work we propose a framework for constructing goodness of fit tests in both low and high-dimensional linear models. We advocate applying regression methods to the scaled residuals following either an ordinary least squares or Lasso fit to the data, and using some proxy for prediction error as the final test statistic. We call this family Residual Prediction (RP) tests. We show that simulation can be used to obtain the critical values for such tests in the low-dimensional setting, and demonstrate using both theoretical results and extensive numerical studies that some form of the parametric bootstrap can do the same when the high-dimensional linear model is under consideration. We show that RP tests can be used to test for significance of groups or individual variables as special cases, and here they compare favourably with state of the art methods, but we also argue that they can be designed to test for as diverse model misspecifications as heteroscedasticity and different types of nonlinearity. 1

Motivation & Objective

  • To develop a unified framework for goodness-of-fit testing applicable to both low- and high-dimensional linear models.
  • To address the challenge of constructing valid hypothesis tests when the number of predictors exceeds the sample size.
  • To enable detection of diverse model misspecifications, including heteroscedasticity and nonlinearity, beyond just variable significance.
  • To provide a computationally feasible method for critical value estimation using simulation or parametric bootstrap.
  • To outperform existing methods in testing significance of individual or grouped variables in high-dimensional settings.

Proposed method

  • Apply ordinary least squares or Lasso to fit the linear model and compute scaled residuals.
  • Use regression of these scaled residuals on the design matrix as a core step to construct test statistics.
  • Employ prediction error from the residual regression as the final test statistic.
  • Use simulation to obtain critical values in low-dimensional settings.
  • Use parametric bootstrap to estimate critical values in high-dimensional settings.
  • Design the test to be sensitive to various model violations, including heteroscedasticity and nonlinearity.

Experimental results

Research questions

  • RQ1Can a unified framework be developed for goodness-of-fit testing in both low- and high-dimensional linear models?
  • RQ2How can critical values for test statistics be reliably computed in high-dimensional settings where traditional methods fail?
  • RQ3To what extent can RP tests detect nonlinearity and heteroscedasticity beyond variable significance testing?
  • RQ4How do RP tests compare in power and size to existing state-of-the-art methods for variable selection?
  • RQ5Can residual regression with prediction error as a test statistic yield valid and powerful inference in high-dimensional models?

Key findings

  • RP tests achieve valid size control and good power for testing individual or group variable significance, outperforming state-of-the-art methods in simulation studies.
  • In high-dimensional settings, the parametric bootstrap provides accurate critical values, enabling reliable inference.
  • The framework is flexible enough to detect diverse model misspecifications, such as heteroscedasticity and nonlinearity, beyond standard significance testing.
  • Simulation studies confirm that critical values obtained via simulation are valid in low-dimensional settings.
  • The use of prediction error from residual regression as a test statistic leads to a robust and interpretable test procedure.
  • The method maintains strong empirical performance across a range of model configurations and error structures.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.