Skip to main content
QUICK REVIEW

[论文解读] How many variables should be entered in a principal component regression equation

Ji Xu, Daniel Hsu|arXiv (Cornell University)|Jun 4, 2019
Statistical Methods and Inference参考文献 3被引用 6
一句话总结

本文研究了在高维设置下主成分回归(PCR)的表现,其中特征为按方差降序排列的不相关高斯分布。结果表明,当特征数 $ p $ 增加时,预测误差呈现出‘双下降’曲线,即使在 $ p > n $ 的情况下也是如此,揭示了模型复杂度与泛化误差之间非单调的关系。

ABSTRACT

We study least squares linear regression over $N$ uncorrelated Gaussian features that are selected in order of decreasing variance. When the number of selected features $p$ is at most the sample size $n$, the estimator under consideration coincides with the principal component regression estimator; when $p>n$, the estimator is the least $\ell_2$ norm solution over the selected features. We give an average-case analysis of the out-of-sample prediction error as $p,n,N o \infty$ with $p/N o \alpha$ and $n/N o \beta$, for some constants $\alpha \in [0,1]$ and $\beta \in (0,1)$. In this average-case setting, the prediction error exhibits a `double descent' shape as a function of $p$.

研究动机与目标

  • 理解主成分回归中特征数 $ p $ 的变化如何影响高维设置下的样本外预测误差。
  • 分析当 $ p > n $ 时预测误差的行为,扩展经典PCR假设的适用范围。
  • 在 $ p, n, N \to \infty $ 且 $ p/N \to \alpha $、$ n/N \to \beta $ 的条件下,刻画平均情况下的预测误差。
  • 研究在随机特征选择与高斯设计下,PCR中是否存在并具有何种性质的‘双下降’现象。

提出的方法

  • 使用 $ N $ 个独立同分布的不相关高斯特征建模回归问题,按方差递减顺序排列。
  • 考虑在前 $ p $ 个特征上进行最小二乘回归,当 $ p \leq n $ 时为PCR估计器,当 $ p > n $ 时为最小 $ \ell_2 $ 范数解。
  • 在 $ N \to \infty $ 的极限下进行渐近分析,其中 $ p/N \to \alpha \in [0,1] $ 且 $ n/N \to \beta \in (0,1) $,假设比值固定。
  • 利用随机矩阵理论和渐近统计分析,推导出作为 $ p $ 函数的平均预测误差。
  • 刻画预测误差随 $ p $ 变化的曲线,显示其具有非单调形状,包含一个最小值和在 $ p > n $ 时的第二次下降。
  • 通过分析高维线性模型中偏差-方差权衡,建立双下降行为的理论基础。

实验结果

研究问题

  • RQ1随着所选特征数 $ p $ 的增加,主成分回归的样本外预测误差如何变化?
  • RQ2即使在不相关高斯特征下,当 $ p > n $ 时预测误差是否仍表现出双下降曲线?
  • RQ3在高维极限下,当 $ p/N \to \alpha $ 且 $ n/N \to \beta $ 时,预测误差的渐近行为如何?
  • RQ4按方差对特征进行排序,如何影响 $ p > n $ 条件下PCR的泛化性能?
  • RQ5在独立同分布高斯特征的平均情况设定下,PCR中的双下降现象能否被解析刻画?

主要发现

  • 预测误差随 $ p $ 变化呈现出‘双下降’形状,最小值出现在 $ p \approx n $ 处,随后在 $ p > n $ 时出现第二次下降。
  • 即使在 $ p > n $ 时,通过所选特征的最小 $ \ell_2 $ 范数解仍能实现比 $ p \leq n $ 更优的泛化性能,这是由于有利的偏差-方差权衡。
  • 在按方差排序的独立同分布高斯特征下,双下降曲线自然出现在平均情况分析中,无需依赖特定信号结构。
  • 渐近预测误差依赖于比值 $ \alpha = p/N $ 和 $ \beta = n/N $,且在某个可能超过 $ n $ 的最优 $ p $ 处达到最小。
  • 分析证实,双下降现象并非过拟合的产物,而是特征选择、维度与正则化之间相互作用的结果。
  • 研究结果将双下降现象的理解从具有随机特征的过参数化模型,扩展到了通过主成分进行结构化特征选择的情形。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。