[论文解读] The out-of-sample $R^2$: estimation and inference
本文提出了一个形式化的样本外 $R^2$ 定义,作为与忽略预测变量的零模型进行模型比较的指标,引入了一种基于数据分割(如交叉验证或自助法)的无偏估计量,并通过德尔塔法推导出标准误。其核心贡献在于实现了对预测性能的统计推断(如置信区间和假设检验),克服了在高维预测场景中长期缺乏对样本外 $R^2$ 的不确定性量化的问题。
Out-of-sample prediction is the acid test of predictive models, yet an independent test dataset is often not available for assessment of the prediction error. For this reason, out-of-sample performance is commonly estimated using data splitting algorithms such as cross-validation or the bootstrap. For quantitative outcomes, the ratio of variance explained to total variance can be summarized by the coefficient of determination or in-sample $R^2$, which is easy to interpret and to compare across different outcome variables. As opposed to the in-sample $R^2$, the out-of-sample $R^2$ has not been well defined and the variability on the out-of-sample $\hat{R}^2$ has been largely ignored. Usually only its point estimate is reported, hampering formal comparison of predictability of different outcome variables. Here we explicitly define the out-of-sample $R^2$ as a comparison of two predictive models, provide an unbiased estimator and exploit recent theoretical advances on uncertainty of data splitting estimates to provide a standard error for the $\hat{R}^2$. The performance of the estimators for the $R^2$ and its standard error are investigated in a simulation study. We demonstrate our new method by constructing confidence intervals and comparing models for prediction of quantitative $ ext{Brassica napus}$ and $ ext{Zea mays}$ phenotypes based on gene expression data.
研究动机与目标
- 解决样本外 $R^2$ 缺乏正式定义和不确定性估计的问题,该指标通常仅以点估计形式报告。
- 通过为样本外 $R^2$ 估计量提供标准误,实现对预测性能的统计推断。
- 在无法获得独立测试数据的场景下,支持模型比较和假设检验(例如,$H_0: R^2 \leq 0$)。
- 通过报告 $R^2$ 的标准误和置信区间,提升预测模型的可重复性和诊断价值。
提出的方法
- 将样本外 $R^2$ 定义为预测模型与忽略预测变量的零模型之间的比较,基于期望的样本外损失。
- 提出一种在交叉验证折或自助样本中对 $R^2$ 进行合并估计的方法,以减少偏差,优于按折平均的方法。
- 通过将德尔塔法应用于均方误差(MSE)和总均方误差(MST)估计量的比值,推导出样本外 $R^2$ 的标准误。
- 利用数据分割程序标准误估计的最新理论进展,确保推断的有效性。
- 采用非参数或刀切法估计 MSE 与 MST 之间的相关性,以提高标准误估计的准确性。
- 通过模拟研究验证该方法,并将其应用于真实组学数据集(拟南芥和玉米,Brassica napus 和 Zea mays),用于构建置信区间和进行模型比较。

实验结果
研究问题
- RQ1能否建立一个形式化且可解释的样本外 $R^2$ 定义,作为模型比较的度量?
- RQ2在使用交叉验证或自助法等数据分割方法时,是否存在样本外 $R^2$ 的无偏估计量?
- RQ3能否为样本外 $R^2$ 推导出可靠的标准误,以支持假设检验和置信区间构建?
- RQ4与基于自助法的替代方法相比,所提出的标准误估计量在控制第一类错误率和覆盖概率方面表现如何?
- RQ5在高维组学数据中,$R^2$ 估计的不确定性在多大程度上影响模型比较和推断?
主要发现
- 在交叉验证折之间对 $R^2$ 进行合并估计的方法几乎无偏,优于按折平均,推荐在交叉验证中使用。
- 德尔塔法可提供有效的样本外 $R^2$ 标准误,支持构建置信区间并检验 $H_0: R^2 \leq 0$。
- 在小样本或高维设置下,标准误估计量存在向上偏倚,但随着预测能力增强,该偏倚会减小。
- 基于自助法的标准误估计量存在向下偏倚,导致推断过于激进,且置信区间覆盖概率低于名义水平。
- 在高维设置下,由于模型拟合的变异性,$R^2$ 估计量的方差较大,提示应避免过度解读 $R^2$ 值的微小差异。
- 报告 $R^2$ 的标准误和置信区间可增强模型诊断能力,支持可重复性,并指导未来研究设计。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。