Skip to main content
QUICK REVIEW

[论文解读] Convex optimization via inertial algorithms with vanishing Tikhonov regularization: fast convergence to the minimum norm solution

Hédy Attouch, Szilárd Csaba László|arXiv (Cornell University)|Apr 24, 2021
Sparse and Compressive Sensing Techniques参考文献 34被引用 8
一句话总结

本论文提出了一种新颖的惯性动态及其一阶算法,结合了消失的Tikhonov正则化与与正则化参数平方根成比例的阻尼,实现了凸优化问题最小范数解的快速、强收敛。该方法在Nesterov加速梯度法的基础上进行了改进,确保了对最优值的快速收敛以及对最小范数最小化器的收敛。

ABSTRACT

In a Hilbertian framework, for the minimization of a general convex differentiable function $f$, we introduce new inertial dynamics and algorithms that generate trajectories and iterates that converge fastly towards the minimizer of $f$ with minimum norm. Our study is based on the non-autonomous version of the Polyak heavy ball method, which, at time $t$, is associated with the strongly convex function obtained by adding to $f$ a Tikhonov regularization term with vanishing coefficient $ε(t)$. In this dynamic, the damping coefficient is proportional to the square root of the Tikhonov regularization parameter $ε(t)$. By adjusting the speed of convergence of $ε(t)$ towards zero, we will obtain both rapid convergence towards the infimal value of $f$, and the strong convergence of the trajectories towards the element of minimum norm of the set of minimizers of $f$. In particular, we obtain an improved version of the dynamic of Su-Boyd-Candès for the accelerated gradient method of Nesterov. This study naturally leads to corresponding first-order algorithms obtained by temporal discretization. In the case of a proper lower semicontinuous and convex function $f$, we study the proximal algorithms in detail, and show that they benefit from similar properties.

研究动机与目标

  • 解决在解集非单点集时实现凸优化中最小范数解的快速收敛的挑战。
  • 克服标准加速梯度方法的局限性,这些方法收敛于某个最小化器,但不一定是范数最小的解。
  • 将Tikhonov正则化与惯性动力学相结合,以平衡快速收敛与对最小范数解的强收敛。
  • 构建连续时间动态及其离散时间对应形式,确保在时间离散化下保持理想的收敛性质。
  • 通过Moreau包络和近端算法将框架扩展至非光滑凸函数,同时保持收敛保证。

提出的方法

  • 提出一种非自治二阶动态(TRIGS),其中Tikhonov正则化参数ε(t)和阻尼系数δ√ε(t)在t → ∞时趋于零。
  • 采用基于李雅普诺夫的分析方法,证明当fₜ为μ-强凸函数时,在瞬态阶段目标误差f(x(t)) − min f = O(e^−√μ t)呈指数衰减。
  • 应用时间离散化,推导出一种用于非光滑凸函数的惯性近端算法(IPATRE-NS),使用Moreau包络fₗₐₘb𝒹a。
  • 在外推步骤中引入松弛参数α > 3,以确保离散序列的稳定性和收敛性。
  • 利用近端微积分恒等式,将离散算法用原函数f及其近端映射表示。
  • 通过利用消失Tikhonov正则化的分层最小化性质,确保强收敛至最小范数最小化器。

实验结果

研究问题

  • RQ1具有消失Tikhonov正则化的惯性动力学能否同时实现对最优值的快速收敛和对最小范数解的强收敛?
  • RQ2如何随时间调节阻尼和正则化参数,以在快速收敛与收敛至最小范数最小化器之间取得平衡?
  • RQ3连续时间动态能否被离散化为保持连续系统收敛性质的一阶算法?
  • RQ4对于非光滑凸函数,收敛速率和对最小范数解的强收敛在多大程度上仍然成立?
  • RQ5与Nesterov加速梯度法及其他惯性算法相比,该方法在收敛速度和解质量方面表现如何?

主要发现

  • 所提出的动态(TRIGS)在fₜ为μ-强凸函数的瞬态阶段,确保f(x(t)) − min f = O(e^−√μ t)。
  • 当ε(t) → 0且∫₀^∞ ε(t) dt = ∞时,连续轨迹x(t)强收敛于argmin f的最小范数解。
  • 离散惯性近端算法(IPATRE-NS)在α > 3条件下,实现f(prox_λf(xₖ)) − min f的o(k^−2s)收敛速率,其中s ∈ [1/2, 1)。
  • 序列(xₖ)满足∑ₖ k^{2s−1} (f(prox_λf(xₖ)) − min f) < ∞,表明误差衰减是可求和的。
  • 若序列最终位于或不位于球B(0, ||x*||)内,则迭代序列(xₖ)强收敛于最小范数最小化器x*。
  • 该方法在Nesterov加速梯度法的基础上进行了改进,确保强收敛至最小范数解,而不仅是一个最小化器。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。