[Paper Review] Convex optimization via inertial algorithms with vanishing Tikhonov regularization: fast convergence to the minimum norm solution
This paper proposes a novel inertial dynamic and associated first-order algorithm that combines vanishing Tikhonov regularization with damping proportional to the square root of the regularization parameter, enabling fast, strong convergence to the minimum norm solution of convex optimization problems. The method improves upon Nesterov's accelerated gradient method by ensuring both rapid convergence to the optimal value and convergence to the minimum norm minimizer.
In a Hilbertian framework, for the minimization of a general convex differentiable function $f$, we introduce new inertial dynamics and algorithms that generate trajectories and iterates that converge fastly towards the minimizer of $f$ with minimum norm. Our study is based on the non-autonomous version of the Polyak heavy ball method, which, at time $t$, is associated with the strongly convex function obtained by adding to $f$ a Tikhonov regularization term with vanishing coefficient $ε(t)$. In this dynamic, the damping coefficient is proportional to the square root of the Tikhonov regularization parameter $ε(t)$. By adjusting the speed of convergence of $ε(t)$ towards zero, we will obtain both rapid convergence towards the infimal value of $f$, and the strong convergence of the trajectories towards the element of minimum norm of the set of minimizers of $f$. In particular, we obtain an improved version of the dynamic of Su-Boyd-Candès for the accelerated gradient method of Nesterov. This study naturally leads to corresponding first-order algorithms obtained by temporal discretization. In the case of a proper lower semicontinuous and convex function $f$, we study the proximal algorithms in detail, and show that they benefit from similar properties.
Motivation & Objective
- Address the challenge of achieving fast convergence to the minimum norm solution in convex optimization, particularly when the solution set is non-singleton.
- Overcome limitations of standard accelerated gradient methods, which converge to a minimizer but not necessarily the one of minimum norm.
- Integrate Tikhonov regularization with inertial dynamics to balance fast convergence and strong convergence to the minimum norm solution.
- Develop a continuous-time dynamic and its discrete-time counterpart that preserve desirable convergence properties under time discretization.
- Extend the framework to non-smooth convex functions via Moreau envelope and proximal algorithms, maintaining convergence guarantees.
Proposed method
- Propose a non-autonomous second-order dynamic (TRIGS) where the Tikhonov regularization parameter ε(t) and damping coefficient δ√ε(t) vanish as t → ∞.
- Use a Lyapunov-based analysis to prove exponential decay of the objective error f(x(t)) − min f = O(e^−√μ t) during the transient phase when fₜ is strongly convex.
- Apply temporal discretization to derive an inertial proximal algorithm (IPATRE-NS) for non-smooth convex functions using the Moreau envelope fₗₐₘb𝒹a.
- Introduce a relaxation parameter α > 3 in the extrapolation step to ensure stability and convergence of the discrete sequence.
- Leverage proximal calculus identities to express the discrete algorithm in terms of the original function f and its proximal mapping.
- Ensure strong convergence to the minimum norm minimizer by exploiting the hierarchical minimization property of vanishing Tikhonov regularization.
Experimental results
Research questions
- RQ1Can inertial dynamics with vanishing Tikhonov regularization achieve both fast convergence to the optimal value and strong convergence to the minimum norm solution?
- RQ2How should the damping and regularization parameters be tuned over time to balance rapid convergence and convergence to the minimum norm minimizer?
- RQ3Can the continuous-time dynamic be discretized into a first-order algorithm that preserves the convergence properties of the continuous system?
- RQ4To what extent do the convergence rates and strong convergence to the minimum norm solution hold for non-smooth convex functions?
- RQ5How does the proposed method compare to Nesterov’s accelerated gradient method and other inertial schemes in terms of convergence speed and solution quality?
Key findings
- The proposed dynamic (TRIGS) ensures f(x(t)) − min f = O(e^−√μ t) during the transient phase when fₜ is μ-strongly convex.
- The continuous trajectory x(t) converges strongly to the minimum norm solution of argmin f as t → ∞, provided ε(t) → 0 and ∫₀^∞ ε(t) dt = ∞.
- The discrete inertial proximal algorithm (IPATRE-NS) achieves o(k^−2s) convergence rate for f(prox_λf(xₖ)) − min f, with s ∈ [1/2, 1), under α > 3.
- The sequence (xₖ) satisfies ∑ₖ k^{2s−1} (f(prox_λf(xₖ)) − min f) < ∞, indicating summable error decay.
- The iterates (xₖ) converge strongly to the minimum norm minimizer x* if the sequence eventually lies inside or outside the ball B(0, ||x*||).
- The method improves upon the Nesterov accelerated gradient method by ensuring strong convergence to the minimum norm solution, not just any minimizer.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.