Skip to main content
QUICK REVIEW

[论文解读] Error bounds in estimating the out-of-sample prediction error using leave-one-out cross validation in high-dimensions

Kamiar Rahnama Rad, Wenda Zhou|arXiv (Cornell University)|Mar 3, 2020
Statistical Methods and Inference参考文献 25被引用 6
一句话总结

本文为高维广义线性模型中的留一法交叉验证(LOO)建立了有限样本误差界,其中特征数 $ p $ 可能超过样本量 $ n $。在温和的正则性条件下且无需稀疏性假设,作者证明了LOO与真实泛化预测误差之间的期望平方误差随着 $ n, p \to \infty $ 而收敛于零,即使当 $ p/n \to \infty $ 时亦成立,从而为LOO在高维设置下经验准确性的理论依据提供了支持。

ABSTRACT

We study the problem of out-of-sample risk estimation in the high dimensional regime where both the sample size $n$ and number of features $p$ are large, and $n/p$ can be less than one. Extensive empirical evidence confirms the accuracy of leave-one-out cross validation (LO) for out-of-sample risk estimation. Yet, a unifying theoretical evaluation of the accuracy of LO in high-dimensional problems has remained an open problem. This paper aims to fill this gap for penalized regression in the generalized linear family. With minor assumptions about the data generating process, and without any sparsity assumptions on the regression coefficients, our theoretical analysis obtains finite sample upper bounds on the expected squared error of LO in estimating the out-of-sample error. Our bounds show that the error goes to zero as $n,p ightarrow \infty$, even when the dimension $p$ of the feature vectors is comparable with or greater than the sample size $n$. One technical advantage of the theory is that it can be used to clarify and connect some results from the recent literature on scalable approximate LO.

研究动机与目标

  • 为在 $ p \gg n $ 的高维设置下留一法交叉验证(LOO)的实证准确性提供理论依据。
  • 分析惩罚回归在广义线性模型族中,LOO估计真实泛化预测误差的期望平方误差。
  • 在不假设真实回归系数稀疏性的前提下,推导LOO估计误差的有限样本上界。
  • 通过统一的理论框架,阐明LOO与可扩展近似LO方法之间的联系。
  • 建立在特征数超过样本数时LOO仍保持一致性的条件。

提出的方法

  • 采用高维渐近框架,其中 $ n, p \to \infty $ 且 $ n/p $ 可能小于1,对LOO进行理论分析。
  • 推导LOO估计中期望平方误差 $ \mathbb{E}[|\text{LO} - \text{Err}_{\text{out}}|^2] $ 的有限样本上界。
  • 利用损失函数 $ \ell(y|\bm{x}^\top\bm{\beta}) $ 和正则化项 $ r(\bm{\beta}) $ 的矩界,结合次高斯和次威布尔尾部性质。
  • 应用集中不等式和矩控制技术,以在给定训练数据的条件下控制LOO估计量的方差。
  • 建立设计矩阵和损失函数的条件,以确保对估计误差实现一致控制。
  • 在数据生成过程的温和矩和正则性假设下,证明LOO的一致性,且无需假设 $ \bm{\beta}^* $ 的稀疏性。

实验结果

研究问题

  • RQ1在何种条件下,留一法交叉验证能在高维设置下一致估计泛化预测误差?
  • RQ2当 $ n $ 和 $ p $ 同时趋于无穷大时,特别是当 $ p > n $ 时,LOO的期望平方误差如何变化?
  • RQ3能否在不假设真实回归系数稀疏性的前提下,推导LOO的有限样本误差界?
  • RQ4LOO与可扩展近似LO方法之间在高维模型中存在何种理论联系?
  • RQ5损失函数和正则化项的矩条件如何影响LOO收敛于真实泛化预测误差?

主要发现

  • 当 $ n, p \to \infty $ 时,LOO与真实泛化预测误差之间的期望平方误差收敛于零,即使 $ p/n \to \infty $ 亦成立。
  • 推导出依赖于损失函数和正则化项矩结构的有限样本上界,明确依赖于 $ \lambda $、$ r(\bm{\beta}^*) $ 和 $ p $。
  • 分析无需对真实系数向量 $ \bm{\beta}^* $ 做任何稀疏性假设,使结果适用于密集高维模型。
  • 在对响应和设计分布的矩条件较弱时,界被证明近乎紧致。
  • 该理论框架通过提供准确性的基准,澄清了可扩展近似LO方法的行为。
  • 推导结果证实,即使在非稀疏情形下LOO仍保持一致,支持其在现代高维问题中的鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。