Skip to main content
QUICK REVIEW

[Paper Review] A systematic approach to general higher-order majorization-minimization algorithms for (non)convex optimization

Ion Necoara, Daniela Lupu|arXiv (Cornell University)|Oct 26, 2020
Sparse and Compressive Sensing Techniques4 citations
TL;DR

This paper proposes a general higher-order majorization-minimization (GHOM) framework for solving nonconvex and nonsmooth optimization problems by iteratively minimizing higher-order surrogate functions that upper-bound the objective with controlled error. The key contribution is convergence guarantees with global sublinear rates for convex problems and locally superlinear rates under uniform convexity or the Kurdyka-Łojasiewicz property, with recovery of known bounds for first- and higher-order methods.

ABSTRACT

Majorization-minimization algorithms consist of successively minimizing a sequence of upper bounds of the objective function so that along the iterations the objective function decreases. Such a simple principle allows to solve a large class of optimization problems, even nonconvex and nonsmooth. We propose a general higher-order majorization-minimization algorithmic framework for minimizing an objective function that admits an approximation (surrogate) such that the corresponding error function has a higher-order Lipschitz continuous derivative. We present convergence guarantees for our new method for general optimization problems with (non)convex and/or (non)smooth objective function. For convex (possibly nonsmooth) problems we provide global sublinear convergence rates, while for problems with uniformly convex objective function we obtain locally faster superlinear convergence rates. We also prove global stationary point guarantees for general nonconvex (possibly nonsmooth) problems and under Kurdyka-Lojasiewicz property of the objective function we derive local convergence rates ranging from sublinear to superlinear for our majorization-minimization algorithm. Moreover, for unconstrained nonconvex problems we derive convergence rates in terms of first- and second-order optimality conditions.

Motivation & Objective

  • To develop a unifying algorithmic framework for higher-order majorization-minimization methods applicable to general (non)convex and (non)smooth optimization problems.
  • To establish global convergence rates for convex problems with nonsmooth or uniformly convex objectives.
  • To provide global asymptotic convergence to stationary points and local convergence rates under the Kurdyka-Łojasiewicz (KL) property for nonconvex problems.
  • To derive convergence rates in terms of first- and second-order optimality conditions for unconstrained nonconvex problems.
  • To recover known complexity bounds from existing tensor and first-order methods as special cases of the proposed framework.

Proposed method

  • The method constructs a higher-order surrogate function that majorizes the objective, where the error function has a p-th order Lipschitz continuous derivative.
  • At each iteration, the algorithm minimizes a local Taylor-type model of the objective using p-th order derivatives, augmented with a penalty term based on the Lipschitz constant of the p-th derivative of the error.
  • The surrogate function satisfies three key properties: majorization of the objective, p-th order smoothness of the error, and exact agreement with the objective and its first p derivatives at the current iterate.
  • Convergence is analyzed using tools from nonconvex optimization, including the Kurdyka-Łojasiewicz inequality and uniform convexity assumptions.
  • The framework generalizes first-order methods (e.g., proximal gradient) and higher-order tensor methods by unifying their convergence analysis under a single abstraction.
  • The algorithm is applied to various problem classes, including composite convex, unconstrained nonconvex, and problems with simple constraints, via tailored surrogate constructions.

Experimental results

Research questions

  • RQ1Can a unified higher-order majorization-minimization framework be developed that generalizes existing first- and higher-order methods for (non)convex and (non)smooth problems?
  • RQ2What convergence rates can be guaranteed for the GHOM algorithm on convex problems with nonsmooth or uniformly convex objectives?
  • RQ3Under what conditions does the GHOM algorithm converge globally to stationary points for nonconvex problems?
  • RQ4How do local convergence rates behave under the Kurdyka-Łojasiewicz property, and what is the dependence on the KL parameter?
  • RQ5Can the GHOM framework recover known convergence bounds for specific classes of problems, such as those solved by proximal gradient or tensor methods?

Key findings

  • For convex (possibly nonsmooth) problems, the GHOM algorithm achieves global sublinear convergence rate of order $\mathcal{O}(k^{-p})$ in objective function values.
  • For uniformly convex problems, the GHOM algorithm achieves locally superlinear convergence in both function values and iterates' distance to the optimal point.
  • For general nonconvex (possibly nonsmooth) problems, the algorithm ensures global asymptotic convergence to stationary points, with $\min_{i=1:k} S(x_i) = \mathcal{O}(k^{-\frac{p}{p+1}})$, where $S(x)$ is the subgradient norm.
  • Under the Kurdyka-Łojasiewicz property, local convergence rates range from sublinear to superlinear depending on the KL parameter, with explicit rate characterization.
  • For smooth unconstrained nonconvex problems, the GHOM algorithm achieves sublinear convergence in both first- and second-order optimality conditions.
  • The framework recovers known convergence bounds for $p=1$ (e.g., proximal gradient and Gauss-Newton) and extends them to $p>1$, with new results for previously unanalyzed problem classes.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.