Skip to main content
QUICK REVIEW

[论文解读] Convergence of Cubic Regularization for Nonconvex Optimization under KL Property

Yi Zhou, Zhe Wang|arXiv (Cornell University)|Aug 22, 2018
Sparse and Compressive Sensing Techniques被引用 6
一句话总结

本文在Kurdyka-Łojasiewicz(KŁ)性质这一广泛几何条件下,建立了非凸优化中立方正则化(CR)的渐近收敛速率。该性质参数化为θ ∈ (0,1],推广了以往的假设。研究发现,CR的收敛速度在阶次上快于一阶方法,其收敛速率依KŁ参数θ的不同,范围从次线性到超线性不等。

ABSTRACT

Cubic-regularized Newton's method (CR) is a popular algorithm that guarantees to produce a second-order stationary solution for solving nonconvex optimization problems. However, existing understandings of the convergence rate of CR are conditioned on special types of geometrical properties of the objective function. In this paper, we explore the asymptotic convergence rate of CR by exploiting the ubiquitous Kurdyka-Lojasiewicz (KL) property of nonconvex objective functions. In specific, we characterize the asymptotic convergence rate of various types of optimality measures for CR including function value gap, variable distance gap, gradient norm and least eigenvalue of the Hessian matrix. Our results fully characterize the diverse convergence behaviors of these optimality measures in the full parameter regime of the KL property. Moreover, we show that the obtained asymptotic convergence rates of CR are order-wise faster than those of first-order gradient descent algorithms under the KL property.

研究动机与目标

  • 理解在一般几何条件下,立方正则化(CR)在非凸优化中的收敛行为。
  • 刻画Kurdyka-Łojasiewicz(KŁ)性质(以θ参数化)如何影响CR的收敛速率。
  • 在相同KŁ条件下,比较CR与一阶梯度方法的收敛速度。
  • 将收敛性分析从特殊情形(如梯度支配性或误差界)推广至完整的KŁ正则函数谱系。
  • 提供一个统一的框架,用于分析CR的收敛性,适用于精确与近似两种变体。

提出的方法

  • 将Kurdyka-Łojasiewicz(KŁ)性质作为非凸函数的统一几何条件,其参数为θ ∈ (0,1]。
  • 推导出广义的KŁ误差界,当θ = 1/2时,该误差界退化为先前工作中提出的局部误差界。
  • 通过有界点到集合的距离distΩ(xk)(即到二阶平稳点集合Ω的距离)来分析CR的收敛性。
  • 建立多种最优性度量的收敛速率:目标函数值差距、变量距离差距、梯度范数以及海森矩阵最小特征值。
  • 利用KŁ误差界,根据θ的取值推导出不同的速率区间,表明收敛可为有限、超线性、线性或次线性。
  • 通过稳定性论证,将分析应用于近似CR变体,表明在采样方案下,相同收敛速率依然成立。

实验结果

研究问题

  • RQ1Kurdyka-Łojasiewicz(KŁ)性质如何影响非凸优化中立方正则化(CR)的收敛速率?
  • RQ2CR在不同KŁ参数θ取值下的收敛行为全谱是什么?
  • RQ3在相同KŁ条件下,CR的收敛速率与一阶梯度下降法相比如何?
  • RQ4KŁ性质能否用于推广基于更强假设(如梯度支配性或局部误差界)的先前收敛结果?
  • RQ5CR的近似变体在多大程度上继承了在KŁ性质下建立的收敛速率?

主要发现

  • CR迭代序列收敛到二阶平稳点集合的速率,关键取决于KŁ参数θ:当θ = 1时,实现有限收敛。
  • 当θ ∈ (1/3, 1)时,到解集的距离以超线性速度衰减,速率上界为exp(−(2θ/(1−θ))^k),对大的k成立。
  • 当θ = 1/3时,收敛为线性,满足distΩ(xk) ≤ Θ(exp(−(k−k₀))),表明指数衰减。
  • 当θ ∈ (0, 1/3)时,收敛为次线性,满足distΩ(xk) ≤ Θ((k−k₀)^{−2θ/(1−3θ})), 表明多项式衰减。
  • 在相同KŁ性质下,CR的收敛速率在阶次上快于一阶梯度下降法,尤其在超线性和有限收敛情形下更为显著。
  • 通过采样方法实现的近似CR算法,其结果可推广,因为收敛动力学在这些近似下保持稳定。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。