Skip to main content
QUICK REVIEW

[论文解读] Conditioning of Random Feature Matrices: Double Descent and Generalization Error

Zhijun Chen, Hayden Schaeffer|arXiv (Cornell University)|Oct 21, 2021
Sparse and Compressive Sensing Techniques参考文献 39被引用 6
一句话总结

本文建立了随机特征矩阵条件数的高概率界,表明当复杂度比 $N/m$ 以 $\log^{-1}(N)$ 或 $\log(m)$ 的方式缩放时,即使不施加正则化,条件数仍保持良好。本文将泛化误差中的双 descent 现象与条件数的行为联系起来,推导出最小二乘、最小范数插值和稀疏回归的显式风险界,其缩放特性在 $m$ 和 $N$ 上最优,且与数据维度无关。

ABSTRACT

We provide (high probability) bounds on the condition number of random feature matrices. In particular, we show that if the complexity ratio $\frac{N}{m}$ where $N$ is the number of neurons and $m$ is the number of data samples scales like $\log^{-1}(N)$ or $\log(m)$, then the random feature matrix is well-conditioned. This result holds without the need of regularization and relies on establishing various concentration bounds between dependent components of the random feature matrix. Additionally, we derive bounds on the restricted isometry constant of the random feature matrix. We prove that the risk associated with regression problems using a random feature matrix exhibits the double descent phenomenon and that this is an effect of the double descent behavior of the condition number. The risk bounds include the underparameterized setting using the least squares problem and the overparameterized setting where using either the minimum norm interpolation problem or a sparse regression problem. For the least squares or sparse regression cases, we show that the risk decreases as $m$ and $N$ increase, even in the presence of bounded or random noise. The risk bound matches the optimal scaling in the literature and the constants in our results are explicit and independent of the dimension of the data.

研究动机与目标

  • 理解高维回归设置下随机特征矩阵的条件性。
  • 在无正则化条件下,建立随机特征矩阵极端奇异值和条件数的高概率界。
  • 将泛化误差的双 descent 行为与设计矩阵的条件数联系起来。
  • 通过随机特征推导出最小二乘、最小范数插值和稀疏回归的显式、与维度无关的风险界。
  • 证明即使在有界或随机噪声下,泛化误差仍随 $m$ 和 $N$ 增加而减小。

提出的方法

  • 利用随机矩阵理论,推导随机特征矩阵中依赖分量的浓度界。
  • 分析随机特征映射的格拉姆矩阵,以界住极端奇异值和条件数。
  • 建立随机特征矩阵的受限等距常数(RIC)的界。
  • 利用压缩感知中的鲁棒恢复结果,将逼近误差与稀疏性和特征质量联系起来。
  • 应用高概率不等式,控制条件数与其期望行为的偏离。
  • 通过连接条件数、恢复误差和泛化误差的一系列不等式,推导风险界。

实验结果

研究问题

  • RQ1在复杂度比 $N/m$ 满足何种条件时,随机特征矩阵以高概率保持良好条件性?
  • RQ2随机特征矩阵的条件数如何驱动泛化误差中的双 descent 现象?
  • RQ3能否在无正则化条件下推导出使用随机特征的回归风险界?其最优缩放特性如何?
  • RQ4最小二乘、最小范数插值和稀疏回归的风险界在 $m$、$N$ 和数据维度上的依赖关系有何异同?
  • RQ5激活函数和权重分布在决定条件性和泛化性能方面起什么作用?

主要发现

  • 当 $N/m \asymp \log^{-1}(N)$ 或 $\log(m)$ 时,随机特征矩阵在无正则化条件下以高概率保持良好条件性。
  • 泛化误差中的双 descent 现象直接由设计矩阵条件数的双 descent 行为驱动。
  • 最小二乘和稀疏回归的风险界缩放为 $\mathcal{O}(N^{-1} + m^{-1/2})$,且具有显式、与维度无关的常数。
  • 对于最小范数插值,即使存在有界或随机噪声,风险也随 $m$ 和 $N$ 增加而减小,且无需正则化。
  • 条件数下界确认了在 $N = m$ 时的病态性,解释了泛化误差在插值阈值处的峰值。
  • 理论界与文献中已知的最优缩放一致,且常数显式且与数据维度 $d$ 无关。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。