[论文解读] A Leisurely Look at Versions and Variants of the Cross Validation Estimator
本文对用于估计分类误差率和AUC的多种交叉验证(CV)版本进行了形式化和比较,证明只有留一法、K折和重复K折CV是非冗余且数学上可靠的。它识别出重复K折CV是唯一平滑的估计器,其在经验上同时匹配条件性能和平均性能的准确性,并建议放弃冗余变体,同时呼吁开展全面的对比研究,并对重复K折CV进行严格的方差估计。
Many versions of cross-validation (CV) exist in the literature; and each version though has different variants. All are used interchangeably by many practitioners; yet, without explanation to the connection or difference among them. This article has three contributions. First, it starts by mathematical formalization of these different versions and variants that estimate the error rate and the Area Under the ROC Curve (AUC) of a classification rule, to show the connection and difference among them. Second, we prove some of their properties and prove that many variants are either redundant or not smooth. Hence, we suggest to abandon all redundant versions and variants and only keep the leave-one-out, the $K$-fold, and the repeated $K$-fold. We show that the latter is the only among the three versions that is smooth and hence looks mathematically like estimating the mean performance of the classification rules. However, empirically, for the known phenomenon of weak correlation, which we explain mathematically and experimentally, it estimates both conditional and mean performance almost with the same accuracy. Third, we conclude the article with suggesting two research points that may answer the remaining question of whether we can come up with a finalist among the three estimators: (1) a comparative study, that is much more comprehensive than those available in literature and conclude no overall winner, is needed to consider a wide range of distributions, datasets, and classifiers including complex ones obtained via the recent deep learning approach. (2) we sketch the path of deriving a rigorous method for estimating the variance of the only smooth version, repeated $K$-fold CV, rather than those ad-hoc methods available in the literature that ignore the covariance structure among the folds of CV.
研究动机与目标
- 对用于误差率和AUC估计的多种交叉验证版本及其变体之间的数学形式化和关联性进行澄清。
- 证明许多现有CV变体是冗余或不平滑的,从而为从标准实践中淘汰它们提供依据。
- 确立重复K折CV是三种主要版本中唯一平滑的估计器,使其能够以数学上一致的方式估计分类规则的平均性能。
- 提出两个关键研究方向:(1) 在多样化数据集、分类器(包括深度学习模型)和分布上开展全面的对比研究;(2) 推导一种严格的方法来估计重复K折CV的方差,该方法需考虑折间协方差结构。
提出的方法
- 本文对不同CV版本(包括留一法、K折和重复K折CV)下的误差率和AUC估计进行了数学形式化。
- 通过数学分析证明,许多CV变体是冗余或不平滑的,表明平滑性对于可靠估计至关重要。
- 证明重复K折CV是唯一平滑的版本,即其行为类似于分类规则平均性能的估计器。
- 通过实证结果表明,尽管CV中存在折间弱相关性,重复K折CV在估计条件性能和平均性能时具有几乎相同的准确性。
- 提出通过考虑折间协方差结构,为重复K折CV推导严格方差估计器的路径。
- 主张放弃非平滑且冗余的CV变体,转而采用三种核心版本:留一法、K折和重复K折CV。
实验结果
研究问题
- RQ1哪些交叉验证版本在数学上是冗余或不平滑的?为何应将其淘汰?
- RQ2尽管各折之间相关性较弱,为何重复K折CV在估计条件性能和平均性能时表现良好?
- RQ3能否通过在多样化数据集、分类器(包括深度学习模型)和分布上的全面对比研究,识别出单一更优的CV估计器?
- RQ4是否可以推导出一种严格估计重复K折CV方差的方法,该方法能正确考虑折级协方差,而非依赖于临时性方法?
- RQ5如何对CV估计器的数学性质进行形式化,以明确其在误差率和AUC估计中相互之间的联系与差异?
主要发现
- 许多现有的交叉验证变体在数学上是冗余或不平滑的,这削弱了其可靠性,因此应从标准实践中淘汰。
- 在三种主要版本(留一法、K折、重复K折)中,只有重复K折CV是平滑的,使其能够以数学上一致的方式估计分类规则的平均性能。
- 尽管各折之间存在弱相关性,重复K折CV在经验上对条件性能和平均性能的估计精度几乎相同。
- 本文结论认为,仅应保留三种CV版本:留一法、K折和重复K折CV,其中重复K折CV因具备平滑性而最具说服力。
- 作者识别出两个关键研究空白:一是需要开展大规模、全面的CV估计器对比研究;二是需要一种能正确考虑重复K折CV中折间协方差的严格方差估计方法。
- 本文拒绝使用临时性方差估计方法,转而倡导一种尊重重复K折CV底层统计结构的方法论框架。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。