[Paper Review] The out-of-sample $R^2$: estimation and inference
This paper proposes a formal definition of the out-of-sample $R^2$ as a model comparison against a null model, introduces an unbiased estimator using data splitting (e.g., cross-validation or bootstrap), and derives a standard error via the delta method. The key contribution is enabling statistical inference—such as confidence intervals and hypothesis tests—on predictive performance, overcoming the long-standing lack of uncertainty quantification for out-of-sample $R^2$ in high-dimensional prediction settings.
Out-of-sample prediction is the acid test of predictive models, yet an independent test dataset is often not available for assessment of the prediction error. For this reason, out-of-sample performance is commonly estimated using data splitting algorithms such as cross-validation or the bootstrap. For quantitative outcomes, the ratio of variance explained to total variance can be summarized by the coefficient of determination or in-sample $R^2$, which is easy to interpret and to compare across different outcome variables. As opposed to the in-sample $R^2$, the out-of-sample $R^2$ has not been well defined and the variability on the out-of-sample $\hat{R}^2$ has been largely ignored. Usually only its point estimate is reported, hampering formal comparison of predictability of different outcome variables. Here we explicitly define the out-of-sample $R^2$ as a comparison of two predictive models, provide an unbiased estimator and exploit recent theoretical advances on uncertainty of data splitting estimates to provide a standard error for the $\hat{R}^2$. The performance of the estimators for the $R^2$ and its standard error are investigated in a simulation study. We demonstrate our new method by constructing confidence intervals and comparing models for prediction of quantitative $ ext{Brassica napus}$ and $ ext{Zea mays}$ phenotypes based on gene expression data.
Motivation & Objective
- To address the lack of a formal definition and uncertainty estimation for out-of-sample $R^2$, which is commonly reported only as a point estimate.
- To enable statistical inference on predictive performance by providing a standard error for the out-of-sample $R^2$ estimator.
- To support model comparison and hypothesis testing (e.g., $H_0: R^2 \leq 0$) in settings where independent test data are unavailable.
- To improve reproducibility and diagnostic value of predictive models by reporting standard errors and confidence intervals for $R^2$.
Proposed method
- Defines the out-of-sample $R^2$ as a comparison between a predictive model and a null model that ignores predictors, using expected out-of-sample loss.
- Proposes a pooling estimator for $R^2$ across cross-validation folds or bootstrap samples to reduce bias, favoring it over fold-wise averaging.
- Derives a standard error for the out-of-sample $R^2$ using the delta method applied to the ratio of mean squared error (MSE) and mean squared total (MST) estimates.
- Utilizes recent theoretical advances on standard error estimation for data splitting procedures to ensure valid inference.
- Employs nonparametric or jackknife estimates of the correlation between $ ext{MSE}$ and $ ext{MST}$ to improve standard error accuracy.
- Validates the method via simulation studies and applies it to real omics datasets (Brassica napus and Zea mays) for confidence interval construction and model comparison.

Experimental results
Research questions
- RQ1Can a formal, interpretable definition of the out-of-sample $R^2$ be established as a model comparison metric?
- RQ2Is there an unbiased estimator for the out-of-sample $R^2$ when using data splitting methods like cross-validation or bootstrap?
- RQ3Can a reliable standard error be derived for the out-of-sample $R^2$ to enable hypothesis testing and confidence interval construction?
- RQ4How does the performance of the proposed standard error estimator compare to bootstrap-based alternatives in terms of type I error control and coverage?
- RQ5To what extent does the uncertainty in $R^2$ estimation affect model comparison and inference in high-dimensional omics data?
Key findings
- The pooling estimator for $R^2$ across cross-validation folds is nearly unbiased, outperforming fold-wise averaging, and is recommended for use with cross-validation.
- The delta method provides a valid standard error for the out-of-sample $R^2$, enabling construction of confidence intervals and testing $H_0: R^2 \leq 0$.
- The standard error estimator is upward biased in small samples or high-dimensional settings, but this bias diminishes as predictive power increases.
- Bootstrap-based standard errors were found to be downward biased, leading to anti-conservative inference and lower-than-nominal confidence interval coverage.
- The variance of the $R^2$ estimator is substantial in high-dimensional settings due to model fitting variability, cautioning against overinterpreting small differences in $R^2$ values.
- Reporting standard errors and confidence intervals for $R^2$ enhances model diagnostics, supports reproducibility, and guides future study design.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.