Skip to main content
QUICK REVIEW

[論文レビュー] The out-of-sample $R^2$: estimation and inference

Stijn Hawinkel, Willem Waegeman|arXiv (Cornell University)|Feb 10, 2023
Statistical Methods and Inference被引用数 4
ひとこと要約

本稿では、予測モデルと予測子を無視する帰無モデルとの比較としての外側サンプル $R^2$ の形式的定義を提示し、データ分割(例:交差検証やブートストラップ)を用いた不偏推定量を導入し、デルタ法を用いて標準誤差を導出する。主な貢献は、予測性能に関する統計的推論(信頼区間や仮説検定など)を可能にし、高次元予測設定における外側サンプル $R^2$ の長年の不確実性評価の欠如を克服することにある。

ABSTRACT

Out-of-sample prediction is the acid test of predictive models, yet an independent test dataset is often not available for assessment of the prediction error. For this reason, out-of-sample performance is commonly estimated using data splitting algorithms such as cross-validation or the bootstrap. For quantitative outcomes, the ratio of variance explained to total variance can be summarized by the coefficient of determination or in-sample $R^2$, which is easy to interpret and to compare across different outcome variables. As opposed to the in-sample $R^2$, the out-of-sample $R^2$ has not been well defined and the variability on the out-of-sample $\hat{R}^2$ has been largely ignored. Usually only its point estimate is reported, hampering formal comparison of predictability of different outcome variables. Here we explicitly define the out-of-sample $R^2$ as a comparison of two predictive models, provide an unbiased estimator and exploit recent theoretical advances on uncertainty of data splitting estimates to provide a standard error for the $\hat{R}^2$. The performance of the estimators for the $R^2$ and its standard error are investigated in a simulation study. We demonstrate our new method by constructing confidence intervals and comparing models for prediction of quantitative $ ext{Brassica napus}$ and $ ext{Zea mays}$ phenotypes based on gene expression data.

研究の動機と目的

  • 外側サンプル $R^2$ に対して形式的定義と不確実性推定の欠如を解消すること。これは、点推定としてのみ報告されることが一般的である。
  • 外側サンプル $R^2$ 推定量の標準誤差を提供することで、予測性能に関する統計的推論を可能にすること。
  • 独立したテストデータが利用できない状況において、モデル比較や仮説検定(例:$H_0: R^2 \leq 0$)を支援すること。
  • 信頼区間や標準誤差の報告を通じて、予測モデルの再現性と診断的価値を向上させること。

提案手法

  • 予測モデルと予測子を無視する帰無モデルとの比較として、期待外側損失を用いて外側サンプル $R^2$ を定義する。
  • 交差検証の各foldやブートストラップ標本における $R^2$ のプーリング推定量を提案し、fold単位の平均化よりも優れていると主張する。
  • 平均二乗誤差(MSE)と平均二乗全変動(MST)推定量の比にデルタ法を適用することで、外側サンプル $R^2$ の標準誤差を導出する。
  • データ分割手順における標準誤差推定に関する最近の理論的進展を活用し、妥当な推論を保証する。
  • 標準誤差の精度を向上させるために、$ ext{MSE}$ と $ ext{MST}$ の間の相関の非パラメトリック推定またはジャックナイフ推定を用いる。
  • シミュレーション研究により手法の妥当性を検証し、実際のオミックスデータセット(ブロスカ・ナプスおよびゼア・マイズ)に適用して信頼区間の構築とモデル比較を実施する。
Figure 2: Diagnostics for the one-dimensional simulation scenario using cross-validation: log10 of the geometric mean of the ratio of estimated to approximated true standard error (SE) of the $R^{2}$ (top panels) and coverage of the confidence intervals (bottom panels) as a function of estimation me
Figure 2: Diagnostics for the one-dimensional simulation scenario using cross-validation: log10 of the geometric mean of the ratio of estimated to approximated true standard error (SE) of the $R^{2}$ (top panels) and coverage of the confidence intervals (bottom panels) as a function of estimation me

実験結果

リサーチクエスチョン

  • RQ1外側サンプル $R^2$ をモデル比較指標として形式的かつ解釈可能な定義を確立できるか?
  • RQ2交差検証やブートストラップなどのデータ分割法を用いる場合に、外側サンプル $R^2$ の不偏推定量が存在するか?
  • RQ3仮説検定や信頼区間の構築を可能にするために、外側サンプル $R^2$ の信頼できる標準誤差を導出できるか?
  • RQ4タイプIエラー制御およびカバレッジの観点から、提案された標準誤差推定量はブートストラップベースの代替手法に比べてどのように性能を発揮するか?
  • RQ5高次元オミックスデータにおける $R^2$ 評価の不確実性が、モデル比較や推論にどの程度影響を及えるか?

主な発見

  • 交差検証の各foldにおける $R^2$ のプーリング推定量はほぼ不偏であり、fold単位の平均化を上回る性能を示し、交差検証と併用する際に推奨される。
  • デルタ法により外側サンプル $R^2$ の有効な標準誤差が得られ、信頼区間の構築や $H_0: R^2 \leq 0$ の仮説検定が可能になる。
  • 小標本または高次元設定では標準誤差推定量が上向きバイアスを示すが、予測力が高くなるにつれてこのバイアスは減少する。
  • ブートストラップベースの標準誤差は下向きバイアスを示し、保守的でない推論と名目水準未満の信頼区間カバレッジを引き起こす。
  • 高次元設定では、モデル適合のばらつきに起因する $R^2$ 推定量の分散が著しく大きくなるため、$R^2$ 値のわずかな差異を過剰に解釈すべきではない。
  • 信頼区間や標準誤差の報告により、モデルの診断的価値が向上し、再現性が高まり、将来的な研究設計の支援が可能になる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。