Skip to main content
QUICK REVIEW

[论文解读] Bagging in overparameterized learning: Risk characterization and risk monotonization

Pratik Patil, Jin‐Hong Du|arXiv (Cornell University)|Oct 20, 2022
Statistical Methods and Inference被引用 4
一句话总结

本文提出一种交叉验证框架,用于在过参数化设置下优化集成预测器,刻画在比例渐近下子样本聚合(subagging)与分块聚合(splagging)变体的渐近预测风险。证明了交叉验证的集成方法可单调化风险曲线,消除双重下降行为,同时在泛化性能上超越全数据训练。

ABSTRACT

Bagging is a commonly used ensemble technique in statistics and machine learning to improve the performance of prediction procedures. In this paper, we study the prediction risk of variants of bagged predictors under the proportional asymptotics regime, in which the ratio of the number of features to the number of observations converges to a constant. Specifically, we propose a general strategy to analyze the prediction risk under squared error loss of bagged predictors using classical results on simple random sampling. Specializing the strategy, we derive the exact asymptotic risk of the bagged ridge and ridgeless predictors with an arbitrary number of bags under a well-specified linear model with arbitrary feature covariance matrices and signal vectors. Furthermore, we prescribe a generic cross-validation procedure to select the optimal subsample size for bagging and discuss its utility to eliminate the non-monotonic behavior of the limiting risk in the sample size (i.e., double or multiple descents). In demonstrating the proposed procedure for bagged ridge and ridgeless predictors, we thoroughly investigate the oracle properties of the optimal subsample size and provide an in-depth comparison between different bagging variants.

研究动机与目标

  • 在比例渐近下,刻画集成岭回归与无岭回归估计器的渐近预测风险。
  • 开发一种通用的交叉验证程序,以选择最优子样本大小,从而单调化极限风险曲线。
  • 比较子样本聚合(有放回)与分块聚合(无放回)在风险表现与单调性方面的差异。
  • 为所提出的框架建立风险单调化的理论保证。
  • 展示最优子样本大小在不同集成变体中所具备的Oracle性质。

提出的方法

  • 利用简单随机抽样中的经典结果,推导集成预测器的精确渐近风险表达式。
  • 应用交叉验证,通过在保留数据上最小化估计预测风险,选择最优子样本大小。
  • 分析子样本聚合(有放回)与分块聚合(无放回)作为不同的集成变体。
  • 将风险表达式表示为纵横比 φ = p/n、特征协方差矩阵与信号强度的函数。
  • 采用比例渐近框架,其中 n, p → ∞ 且 p/n → φ。
  • 通过在多种模型(M-ISO-LI, M-AR1-LI)上进行大量模拟,验证理论结果,涵盖不同信噪比(SNR)与特征结构。

实验结果

研究问题

  • RQ1在过参数化线性模型中,子样本聚合与分块聚合的岭回归与无岭回归估计器的精确渐近预测风险是什么?
  • RQ2能否设计一种交叉验证程序,可严格证明地单调化作为样本大小或纵横比函数的渐近风险?
  • RQ3在风险表现与模型误设的鲁棒性方面,子样本聚合与分块聚合如何比较?
  • RQ4通过交叉验证选择的最优子样本大小在有限样本中是否能实现类似Oracle的性能?
  • RQ5在何种条件下,集成方法可消除泛化误差中的双重下降现象?

主要发现

  • 所提出的交叉验证框架成功单调化了集成预测器的渐近风险,消除了如双重下降等非单调行为。
  • 在高维设置下,采用多袋子的子样本聚合(有放回)始终显著降低预测风险,优于未集成的预测器。
  • 对于无岭与岭回归估计器,通过交叉验证选择的最优子样本大小,其风险表现接近理论Oracle水平。
  • 在相同条件下,分块聚合(无放回)的渐近风险通常低于子样本聚合(有放回),尤其当 M 较大时。
  • 在模拟实验中,交叉验证的集成预测器在所有信噪比水平与特征相关结构下,均优于全数据估计器。
  • 交叉验证集成预测器的风险曲线在纵横比 φ 上保持单调,与理论预测一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。