Skip to main content
QUICK REVIEW

[Paper Review] Iteration complexity analysis of random coordinate descent methods for $\ell_0$ regularized convex problems

Andrei Pătraşcu, Ion Necoara|arXiv (Cornell University)|Mar 26, 2014
Sparse and Compressive Sensing Techniques24 references4 citations
TL;DR

This paper proposes a unified convergence analysis for random block coordinate descent methods applied to $β_0$-regularized convex optimization problems, where the objective combines a smooth convex function and an $β_0$ quasinorm penalty. It establishes almost sure convergence to local minima within restricted classes defined by approximation versions of the objective, and proves linear convergence in probability under strong convexity.

ABSTRACT

In this paper we analyze a family of general random block coordinate descent methods for the minimization of $\ell_0$ regularized optimization problems, i.e. the objective function is composed of a smooth convex function and the $\ell_0$ regularization. Our family of methods covers particular cases such as random block coordinate gradient descent and random proximal coordinate descent methods. We analyze necessary optimality conditions for this nonconvex $\ell_0$ regularized problem and devise a separation of the set of local minima into restricted classes based on approximation versions of the objective function. We provide a unified analysis of the almost sure convergence for this family of block coordinate descent algorithms and prove that, for each approximation version, the limit points are local minima from the corresponding restricted class of local minimizers. Under the strong convexity assumption, we prove linear convergence in probability for our family of methods.

Motivation & Objective

  • To address the challenge of solving nonconvex $β_0$-regularized optimization problems that are combinatorially hard and lack exact equivalence between regularized and sparsity-constrained formulations.
  • To develop a unified convergence framework for random block coordinate descent methods tailored to $β_0$-regularized problems, covering both gradient and proximal variants.
  • To classify local minima into restricted classes based on approximation versions of the objective function, enabling finer characterization of convergence behavior.
  • To establish almost sure convergence of the algorithmic iterates to local minima within the corresponding restricted class for each approximation version.
  • To prove linear convergence in probability under strong convexity assumptions, providing iteration complexity guarantees for the proposed methods.

Proposed method

  • Proposes a family of random block coordinate descent algorithms that minimize an approximation of the $β_0$-regularized objective function, updating one block of variables at a time while fixing others.
  • Introduces a separation of the set of local minima into restricted classes based on the choice of approximation version of the objective function.
  • Uses a general algorithmic framework that subsumes specific methods such as random block coordinate gradient descent and random proximal coordinate descent.
  • Employs a stochastic index selection rule for block updates, ensuring random exploration of variable blocks at each iteration.
  • Analyzes convergence using optimality conditions tailored to the nonconvex $β_0$-regularized problem, distinguishing between different types of local minimizers.
  • Establishes convergence under strong convexity by proving linear convergence in probability, leveraging the structure of the approximation and the randomness in block selection.

Experimental results

Research questions

  • RQ1What are the necessary optimality conditions for $β_0$-regularized nonconvex optimization problems, and how can local minima be classified?
  • RQ2How does the choice of approximation version of the objective function affect the convergence behavior of random block coordinate descent methods?
  • RQ3Can a unified convergence analysis be established for random block coordinate descent methods applied to $β_0$-regularized problems?
  • RQ4Under what conditions does the proposed method achieve linear convergence in probability?
  • RQ5How do the proposed algorithms compare in practice to existing methods like IHTA in terms of convergence speed and sparsity of solutions?

Key findings

  • The proposed random block coordinate descent methods converge almost surely to local minima within restricted classes defined by the approximation version of the objective function.
  • For strongly convex problems, the method achieves linear convergence in probability, indicating favorable iteration complexity under this condition.
  • Algorithm (RCD-IHT-$u^e$) required significantly fewer full iterations than (IHTA) to reach superior function values, especially as problem dimension increased.
  • In experiments on $β_2$-regularized logistic loss models, (RCD-IHT-$u^e$) achieved the best function values with fewer full iterations, demonstrating scalability and efficiency.
  • The method maintained high sparsity in solutions while achieving low objective function values, with sparsity levels ranging from 15 to 767 nonzero components across test instances.
  • Performance of (RCD-IHT-$u^e$) scaled well with increasing problem size, achieving convergence in as few as 12 full iterations even for problems with 2500 variables.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.