Skip to main content
QUICK REVIEW

[论文解读] A Risk Ratio Comparison of $l_0$ and $l_1$ Penalized Regression

Kory D. Johnson, Dongyu Lin|arXiv (Cornell University)|Oct 21, 2015
Sparse and Compressive Sensing Techniques参考文献 25被引用 15
一句话总结

本文比较了 $l_0$ 和 $l_1$ 惩罚回归在特征选择中的表现,表明尽管 $l_1$ 正则化由于凸性具有计算效率优势,但其预测性能可能远差于 $l_0$ 正则化。主要贡献在于证明:通过逐步回归获得的 $l_0$ 近似解通常优于精确的 $l_1$ 解,且使用 $l_0$ 准则选择最佳 $l_1$ 模型可显著提升准确性。

ABSTRACT

There has been an explosion of interest in using $l_1$-regularization in place of $l_0$-regularization for feature selection. We present theoretical results showing that while $l_1$-penalized linear regression never outperforms $l_0$-regularization by more than a constant factor, in some cases using an $l_1$ penalty is infinitely worse than using an $l_0$ penalty. We also show that the "optimal" $l_1$ solutions are often inferior to $l_0$ solutions found using stepwise regression. We also compare algorithms for solving these two problems and show that although solutions can be found efficiently for the $l_1$ problem, the "optimal" $l_1$ solutions are often inferior to $l_0$ solutions found using greedy classic stepwise regression. Furthermore, we show that solutions obtained by solving the convex $l_1$ problem can be improved by selecting the best of the $l_1$ models (for different regularization penalties) by using an $l_0$ criterion. In other words, an approximate solution to the right problem can be better than the exact solution to the wrong problem.

研究动机与目标

  • 正式分析并比较高维特征选择中 $l_0$ 与 $l_1$ 惩罚回归的预测风险性能。
  • 探究为何尽管 $l_1$ 正则化具有计算优势,但在实践中常表现不如 $l_0$ 方法。
  • 评估最优 $l_1$ 解是否确实优于通过贪心算法(如逐步回归)获得的 $l_0$ 解。
  • 提出一种混合方法,利用 $l_1$ 优化生成候选模型,并通过 $l_0$ 准则选择最优模型。
  • 证明:对正确问题($l_0$)的近似解,可能优于对错误问题($l_1$)的精确解。

提出的方法

  • 在独立同分布高斯噪声的正态线性模型下,对 $l_0$ 与 $l_1$ 估计器的极小极大风险比进行理论分析。
  • 构建合成数据示例,显示 $l_1$ 正则化相比 $l_0$ 可导致任意差的预测性能。
  • 在源自 NP-难子集选择问题的精确覆盖问题上,对 Lasso($l_1$)与前向逐步回归(基于 $l_0$)进行经验比较。
  • 使用 LARS 算法生成 $l_1$ 正则化路径,并基于 $l_0$ 惩罚的训练误差选择最佳模型。
  • 使用合成数据与真实世界数据集,评估模型稀疏性与预测风险,风险以未来观测的期望预测误差衡量。
  • 应用次模优化原理,解释贪心 $l_0$-基搜索在高维设置下的有效性。

实验结果

研究问题

  • RQ1在最坏情况下,$l_1$ 正则化回归的预测风险与 $l_0$ 正则化回归相比如何?
  • RQ2$l_1$ 正则化在预测精度上是否可能远差于 $l_0$ 正则化?
  • RQ3最优 $l_1$ 解是否总是优于通过贪心方法(如逐步回归)获得的 $l_0$ 解?
  • RQ4能否通过使用 $l_0$ 准则从 $l_1$ 模型中选择来提升 $l_1$ 方法的性能?
  • RQ5在何种条件下,$l_0$ 的近似解会优于 $l_1$ 的精确解?

主要发现

  • 在最坏情况下,$l_1$ 与 $l_0$ 的风险比可呈二次增长并趋于无穷大,表明 $l_1$ 可能无限差于 $l_0$。
  • 在稀疏系统中,风险比的上确界 $\sup_{\boldsymbol{\beta}} \frac{R(\boldsymbol{\beta}, \hat{\boldsymbol{\beta}}_{l_1})}{R(\boldsymbol{\beta}, \hat{\boldsymbol{\beta}}_{l_0})}$ 趋于无穷,说明 $l_1$ 可能表现极差。
  • 前向逐步回归(一种贪心 $l_0$-基方法)始终生成比 Lasso 更稀疏的模型,且预测风险更低。
  • 在精确覆盖问题上,逐步回归选择的子集更少(例如,$n=30$ 时为 10 个 vs. 29 个),且平方误差和更低,表明模型选择更优。
  • 通过最小化 $l_0$-惩罚训练误差选择的最优 $l_1$ 模型,始终优于最优 $l_1$ 解,证明 $l_0$ 准则是更有效的模型评估标准。
  • 对正确问题($l_0$)的近似解,可能优于对错误问题($l_1$)的精确解,凸显问题建模的重要性,而非单纯追求计算上的精确。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。