Skip to main content
QUICK REVIEW

[论文解读] Asymptotic Risk of Least Squares Minimum Norm Estimator under the Spike Covariance Model.

Yasaman Mahdaviyeh, Zacharie Naulet|arXiv (Cornell University)|Dec 31, 2019
Random Matrices and Applications参考文献 12被引用 5
一句话总结

本文分析了在高维线性回归中,当参数数量 d 的增长速度快于样本量 n 时,最小二乘最小范数估计量的渐近风险。在具有固定数量发散特征值的尖峰协方差模型下,风险收敛于零,表明即使在高维情况下,通过最小范数估计进行插值也能实现低渐近风险。

ABSTRACT

One of the recent approaches to explain good performance of neural networks has focused on their ability to fit training data perfectly (interpolate) without overfitting. It has been shown that this is not unique to neural nets, and that it happens with simpler models such as kernel regression, too Belkin et al. (2018b); Tengyuan Liang (2018). Consequently, there has been quite a few works that give conditions for low risk or optimality of interpolating models, see for example Belkin et al. (2018a, 2019b). One of the simpler models where interpolation has been studied recently is least squares solution for linear regression. In this case, interpolation is only guaranteed to happen in high dimensional setting where the number of parameters exceeds number of samples; therefore, least squares solution is not necessarily unique. However, minimum norm solution is unique, can be written in closed form, and gradient descent starting at the origin converges to it (Hastie et al., 2019). This has, at least partially, motivated several works that study risk of minimum norm least squares estimator for linear regression. Continuing in a similar vein, we study the asymptotic risk of minimum norm least squares estimator when number of parameters $d$ depends on $n$, and $\frac{d}{n} ightarrow \infty$. In this high dimensional setting, to make inference feasible, it is usually assumed that true parameters or data have some underlying low dimensional structure such as sparsity, or vanishing eigenvalues of population covariance matrix. Here, we restrict ourselves to spike covariance matrices, where a fixed finite number of eigenvalues grow with $n$ and are much larger than the rest of the eigenvalues, which are (asymptotically) in the same order. We show that in this setting the risk can vanish.

研究动机与目标

  • 理解在 d/n → ∞ 的高维线性回归中,最小范数最小二乘估计量的风险行为。
  • 研究尽管存在过参数化,通过最小范数解进行插值是否仍能实现低风险。
  • 在结构化协方差模型——特别是具有有限个发散特征值的尖峰协方差模型——下分析渐近风险。
  • 建立最小范数估计量风险在高维设定下趋于消失的条件。

提出的方法

  • 将最小二乘最小范数估计量形式化为在拟合训练数据的前提下最小化 L2 范数的约束优化问题的解。
  • 假设高维渐近框架,其中 d/n → ∞,且 d 和 n 同时趋于无穷。
  • 对设计矩阵施加尖峰协方差模型,其中固定数量的特征值随 n 增长,其余特征值有界且渐近阶相同。
  • 使用随机矩阵理论工具推导估计量的渐近风险,重点关注在指定谱结构下估计量的行为。
  • 在尖峰模型下分析当 n → ∞ 且 d/n → ∞ 时风险表达式的极限行为。
  • 证明在尖峰模型下,风险收敛于零,表明在此框架中通过最小范数估计进行插值可以保持一致性。

实验结果

研究问题

  • RQ1当 d/n → ∞ 时,最小范数最小二乘估计量的渐近风险是什么?
  • RQ2在何种关于总体协方差矩阵的条件下,最小范数估计量的风险会消失?
  • RQ3尖峰协方差结构——以固定数量的发散特征值为特征——如何影响插值估计量的风险行为?
  • RQ4在过参数化设置的高维环境中,通过最小范数解进行插值是否能实现低风险?
  • RQ5是否存在低维结构(通过尖峰模型体现)可使估计在参数数量超过样本数量时仍保持一致?

主要发现

  • 在 d/n → ∞ 的高维设定下,最小范数最小二乘估计量在尖峰协方差模型下的渐近风险收敛于零。
  • 风险消失由尖峰结构驱动:固定有限数量的特征值随 n 增长,其余特征值保持有界且渐近阶相同。
  • 该结果表明,即使在插值不可避免的过参数化设置中,只要设计协方差具有低维尖峰结构,最小范数解仍能实现低风险。
  • 该分析证实,只要底层协方差具有尖峰结构,最小范数解在高维设定下并不必然导致高风险。
  • 研究结果支持插值估计量中“良性过拟合”现象的普遍性,表明对设计结构(如尖峰协方差)的假设可确保低风险。
  • 该结果为梯度下降在过参数化模型中经验成功的理论基础提供了支持,因为其收敛于最小范数解,而在此设定下该解的风险趋于消失。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。