Skip to main content
QUICK REVIEW

[Paper Review] The out-of-sample $R^2$: estimation and inference

Stijn Hawinkel, Willem Waegeman|arXiv (Cornell University)|Feb 10, 2023
Statistical Methods and Inference4 citations
TL;DR

This paper proposes a formal definition of the out-of-sample $R^2$ as a model comparison against a null model, introduces an unbiased estimator using data splitting (e.g., cross-validation or bootstrap), and derives a standard error via the delta method. The key contribution is enabling statistical inference—such as confidence intervals and hypothesis tests—on predictive performance, overcoming the long-standing lack of uncertainty quantification for out-of-sample $R^2$ in high-dimensional prediction settings.

ABSTRACT

Out-of-sample prediction is the acid test of predictive models, yet an independent test dataset is often not available for assessment of the prediction error. For this reason, out-of-sample performance is commonly estimated using data splitting algorithms such as cross-validation or the bootstrap. For quantitative outcomes, the ratio of variance explained to total variance can be summarized by the coefficient of determination or in-sample $R^2$, which is easy to interpret and to compare across different outcome variables. As opposed to the in-sample $R^2$, the out-of-sample $R^2$ has not been well defined and the variability on the out-of-sample $\hat{R}^2$ has been largely ignored. Usually only its point estimate is reported, hampering formal comparison of predictability of different outcome variables. Here we explicitly define the out-of-sample $R^2$ as a comparison of two predictive models, provide an unbiased estimator and exploit recent theoretical advances on uncertainty of data splitting estimates to provide a standard error for the $\hat{R}^2$. The performance of the estimators for the $R^2$ and its standard error are investigated in a simulation study. We demonstrate our new method by constructing confidence intervals and comparing models for prediction of quantitative $ ext{Brassica napus}$ and $ ext{Zea mays}$ phenotypes based on gene expression data.

Motivation & Objective

  • To address the lack of a formal definition and uncertainty estimation for out-of-sample $R^2$, which is commonly reported only as a point estimate.
  • To enable statistical inference on predictive performance by providing a standard error for the out-of-sample $R^2$ estimator.
  • To support model comparison and hypothesis testing (e.g., $H_0: R^2 \leq 0$) in settings where independent test data are unavailable.
  • To improve reproducibility and diagnostic value of predictive models by reporting standard errors and confidence intervals for $R^2$.

Proposed method

  • Defines the out-of-sample $R^2$ as a comparison between a predictive model and a null model that ignores predictors, using expected out-of-sample loss.
  • Proposes a pooling estimator for $R^2$ across cross-validation folds or bootstrap samples to reduce bias, favoring it over fold-wise averaging.
  • Derives a standard error for the out-of-sample $R^2$ using the delta method applied to the ratio of mean squared error (MSE) and mean squared total (MST) estimates.
  • Utilizes recent theoretical advances on standard error estimation for data splitting procedures to ensure valid inference.
  • Employs nonparametric or jackknife estimates of the correlation between $ ext{MSE}$ and $ ext{MST}$ to improve standard error accuracy.
  • Validates the method via simulation studies and applies it to real omics datasets (Brassica napus and Zea mays) for confidence interval construction and model comparison.
Figure 2: Diagnostics for the one-dimensional simulation scenario using cross-validation: log10 of the geometric mean of the ratio of estimated to approximated true standard error (SE) of the $R^{2}$ (top panels) and coverage of the confidence intervals (bottom panels) as a function of estimation me
Figure 2: Diagnostics for the one-dimensional simulation scenario using cross-validation: log10 of the geometric mean of the ratio of estimated to approximated true standard error (SE) of the $R^{2}$ (top panels) and coverage of the confidence intervals (bottom panels) as a function of estimation me

Experimental results

Research questions

  • RQ1Can a formal, interpretable definition of the out-of-sample $R^2$ be established as a model comparison metric?
  • RQ2Is there an unbiased estimator for the out-of-sample $R^2$ when using data splitting methods like cross-validation or bootstrap?
  • RQ3Can a reliable standard error be derived for the out-of-sample $R^2$ to enable hypothesis testing and confidence interval construction?
  • RQ4How does the performance of the proposed standard error estimator compare to bootstrap-based alternatives in terms of type I error control and coverage?
  • RQ5To what extent does the uncertainty in $R^2$ estimation affect model comparison and inference in high-dimensional omics data?

Key findings

  • The pooling estimator for $R^2$ across cross-validation folds is nearly unbiased, outperforming fold-wise averaging, and is recommended for use with cross-validation.
  • The delta method provides a valid standard error for the out-of-sample $R^2$, enabling construction of confidence intervals and testing $H_0: R^2 \leq 0$.
  • The standard error estimator is upward biased in small samples or high-dimensional settings, but this bias diminishes as predictive power increases.
  • Bootstrap-based standard errors were found to be downward biased, leading to anti-conservative inference and lower-than-nominal confidence interval coverage.
  • The variance of the $R^2$ estimator is substantial in high-dimensional settings due to model fitting variability, cautioning against overinterpreting small differences in $R^2$ values.
  • Reporting standard errors and confidence intervals for $R^2$ enhances model diagnostics, supports reproducibility, and guides future study design.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.