Skip to main content
QUICK REVIEW

[论文解读] Accelerated Stochastic Subgradient Methods under Local Error Bound Condition

Yi Xu, Qihang Lin|arXiv (Cornell University)|Jul 4, 2016
Sparse and Compressive Sensing Techniques参考文献 4被引用 10
一句话总结

本文通过利用局部误差界条件,提出两种加速的随机次梯度方法,用于非强凸优化。通过在历史解的缩小局部区域内迭代求解问题——采用显式球约束或隐式正则化——实现了对初始距离最优集的对数依赖性改进的高概率迭代复杂度,包括在 $ε$-正则化多面体损失问题中 $ε$-最优解的对数复杂度。

ABSTRACT

In this paper, we propose two {\bf accelerated stochastic subgradient} methods for stochastic non-strongly convex optimization problems by leveraging a generic local error bound condition. The novelty of the proposed methods lies at smartly leveraging the recent historical solution to tackle the variance in the stochastic subgradient. The key idea of both methods is to iteratively solve the original problem approximately in a local region around a recent historical solution with size of the local region gradually decreasing as the solution approaches the optimal set. The difference of the two methods lies at how to construct the local region. The first method uses an explicit ball constraint and the second method uses an implicit regularization approach. For both methods, we establish the improved iteration complexity in a high probability for achieving an $\epsilon$-optimal solution. Besides the improved order of iteration complexity with a high probability, the proposed algorithms also enjoy a logarithmic dependence on the distance of the initial solution to the optimal set. We also consider applications in machine learning and demonstrate that the proposed algorithms enjoy faster convergence than the traditional stochastic subgradient method. For example, when applied to the $\ell_1$ regularized polyhedral loss minimization (e.g., hinge loss, absolute loss), the proposed stochastic methods have a logarithmic iteration complexity.

研究动机与目标

  • 解决传统随机次梯度方法在非强凸优化中收敛缓慢的问题。
  • 在局部误差界条件下,改进实现 $ε$-最优解的迭代复杂度。
  • 通过利用历史解信息,减少对初始距离最优集的依赖。
  • 设计在处理随机次梯度方差的同时保持加速特性的方法。
  • 在机器学习应用中展示实际优势,如 $µ_1$-正则化合页损失和绝对损失最小化。

提出的方法

  • 该方法在最近历史解的局部区域内进行,随着迭代接近最优集,区域大小逐渐减小。
  • 第一种方法在历史解处施加显式的球约束,以限制搜索空间。
  • 第二种方法采用隐式正则化方法,以塑造局部区域而不使用显式约束。
  • 两种方法均在不断缩小的局部区域内近似最小化原始目标函数,以减少次梯度方差。
  • 算法设计旨在实现改进的迭代复杂度的高概率收敛保证。
  • 利用局部误差界条件,确保当解接近最优集时,进展速度加快。

实验结果

研究问题

  • RQ1在不依赖强凸性的情况下,能否在局部误差界条件下加速随机次梯度方法?
  • RQ2在显式球约束与隐式正则化之间进行选择,如何影响收敛性能?
  • RQ3所提方法在实现 $ε$-最优解时的高概率迭代复杂度是多少?
  • RQ4所提方法是否实现了对初始距离最优集的对数依赖性?
  • RQ5这些方法能否有效应用于如 $µ_1$-正则化多面体损失最小化等机器学习问题?

主要发现

  • 所提方法相比标准随机次梯度方法,实现了改进的高概率迭代复杂度。
  • 迭代复杂度对初始距离最优集表现出对数依赖性,相比标准方法中的多项式依赖,这是显著改进。
  • 对于 $µ_1$-正则化多面体损失最小化问题(例如合页损失、绝对损失),迭代复杂度在所需精度 $ε$ 上为对数关系。
  • 在局部子问题中使用历史解,有效降低了随机次梯度的方差。
  • 隐式正则化变体在性能上与显式球约束方法相当,同时提供了更高的算法灵活性。
  • 在机器学习应用中的实证结果表明,收敛速度优于传统随机次梯度方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。