Skip to main content
QUICK REVIEW

[論文レビュー] Optimality of Training/Test Size and Resampling Effectiveness of Cross-Validation Estimators of the Generalization Error

Georgios Afendras, Marianthi Markatou|arXiv (Cornell University)|Nov 10, 2015
Statistical Methods and Inference参考文献 23被引用数 12
ひとこと要約

本稿では、広範な損失関数のクラスに対して、k-fold交差検証における最適な訓練データサイズが、合計サンプルサイズのちょうど半分であることを確立している。一般化誤差推定の分散を最小化する。本稿ではリサンプリング効果性を指標として導入し、分散最小化を用いて最適な分割ルールを導出し、論理的および実験的にロジスティック回帰を用いた分類において検証された。

ABSTRACT

An important question in constructing Cross Validation (CV) estimators of the generalization error is whether rules can be established that allow "optimal" selection of the size of the training set, for fixed sample size $n$. We define the {\it resampling effectiveness} of random CV estimators of the generalization error as the ratio of the limiting value of the variance of the CV estimator over the estimated from the data variance. The variance and the covariance of different average test set errors are independent of their indices, thus, the resampling effectiveness depends on the correlation and the number of repetitions used in the random CV estimator. We discuss statistical rules to define optimality and obtain the "optimal" training sample size as the solution of an appropriately formulated optimization problem. We show that in a broad class of loss functions the optimal training size equals half of the total sample size, independently of the data distribution. We optimally select the number of folds in $k$-fold cross validation and offer a computational procedure for obtaining the optimal splitting in the case of classification (via logistic regression). We substantiate our claims both, theoretically and empirically.

研究の動機と目的

  • 固定された合計サンプルサイズ下で、一般化誤差推定の分散を最小化する交差検証における最適な訓練データサイズを特定すること。
  • リサンプリング効果性を、ランダムな交差検証における限界分散と実効的分散の比として定義・定量化すること。
  • 固定された合計サンプルサイズ下で、最適な訓練/テストサイズ選択のための最適化問題を定式化し、解くこと。
  • k-fold交差検証に結果を拡張し、分類タスクにおけるロジスティック回帰を用いた計算的手順を提供すること。
  • 理論的分析と実験的シミュレーションを通じて、理論的考察を検証すること。

提案手法

  • リサンプリング効果性を、交差検証推定量の限界分散をデータからの推定分散で割った比として定義する。
  • 予測誤差の漸近的正規性を用い、折りたたみと繰り返しの間で予測誤差の同時モーメントを導出する。
  • 予測誤差のベクトルに対する多変量正規近似を適用し、その共分散構造は訓練サイズと折りたたみの重複に依存する。
  • データポイントが訓練セットに含まれるかどうかの指標関数の期待値を用いて、交差検証推定量の分散を訓練サイズ $ n_1 $ の関数として導出する。
  • この分散を最小化する最適化問題を解き、広範な損失関数のクラスに対して最小値が $ n_1 = n/2 $ で達成されることを示す。
  • 推定分散を最小化する基準に基づき、ロジスティック回帰を用いた分類タスクにおける最適な分割のための計算アルゴリズムを提案する。

実験結果

リサーチクエスチョン

  • RQ1固定された合計サンプルサイズ下で、一般化誤差推定の分散を最小化する訓練データサイズは何か?
  • RQ2ランダムな交差検証において、リサンプリング効果性を形式的に定義・定量化する方法は何か?
  • RQ3最適な訓練サイズはデータ分布や損失関数に依存するか?その条件は何か?
  • RQ4k-折り交差検証における最適な折りたたみ数は、分散最小化から導出可能か?
  • RQ5本手法は、ロジスティック回帰を用いた分類タスクにおいて、実験的にどの程度の性能を示すか?

主な発見

  • 広範な損失関数のクラスに対して、最適な訓練データサイズは、母集団のデータ分布に依存せず、正確に $ n/2 $ である。
  • リサンプリング効果性は、テストセット誤差間の相関とランダムな交差検証における繰り返し回数に依存し、相関が高いほど効果性が低下する。
  • 漸近的分布的分析を通じて、交差検証推定量の分散が訓練データサイズが合計サンプルサイズの半分であるとき最小化されることを示した。
  • 最適な折りたたみ数は、同じ分散最小化フレームワークから導出可能であり、理論的裏付けがある。
  • 推定分散を最小化する基準に基づき、ロジスティック回帰を用いた分類タスクにおける最適な分割選択のための計算手順を提供した。
  • 理論的結果は実験的シミュレーションによって裏付けられ、$ n_1 = n/2 $ がさまざまな設定で最適であることが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。