[Paper Review] On the Linear Convergence of the Alternating Direction Method of Multipliers
This paper establishes the global linear convergence of the Alternating Direction Method of Multipliers (ADMM) for minimizing the sum of any number of convex separable functions, even without strong convexity. The analysis relies on error bounds and proximal residuals, proving linear convergence under a sufficiently small dual stepsize, which resolves a long-standing open question for multi-block and non-strongly convex problems.
We analyze the convergence rate of the alternating direction method of multipliers (ADMM) for minimizing the sum of two or more nonsmooth convex separable functions subject to linear constraints. Previous analysis of the ADMM typically assumes that the objective function is the sum of only two convex functions defined on two separable blocks of variables even though the algorithm works well in numerical experiments for three or more blocks. Moreover, there has been no rate of convergence analysis for the ADMM without strong convexity in the objective function. In this paper we establish the global linear convergence of the ADMM for minimizing the sum of any number of convex separable functions. This result settles a key question regarding the convergence of the ADMM when the number of blocks is more than two or if the strong convexity is absent. It also implies the linear convergence of the ADMM for several contemporary applications including LASSO, Group LASSO and Sparse Group LASSO without any strong convexity assumption. Our proof is based on estimating the distance from a dual feasible solution to the optimal dual solution set by the norm of a certain proximal residual, and by requiring the dual stepsize to be sufficiently small.
Motivation & Objective
- To resolve the open question of whether ADMM converges linearly for problems with more than two blocks or without strong convexity.
- To extend convergence analysis beyond the classical two-block setting, where ADMM is known to converge linearly.
- To provide a theoretical foundation for ADMM's empirical success in applications like LASSO, Group LASSO, and Sparse Group LASSO without strong convexity assumptions.
- To establish linear convergence via error bound techniques and proximal residual estimation, independent of strong convexity.
- To generalize convergence guarantees to a broad class of structured convex optimization problems with multiple separable blocks.
Proposed method
- The authors analyze the ADMM by estimating the distance from a dual feasible solution to the optimal dual solution set using the norm of a proximal residual.
- They introduce a novel error bound condition that links the primal infeasibility and dual infeasibility to the norm of the proximal residual.
- The proof relies on strong convexity of auxiliary functions constructed from the augmented Lagrangian, even when the original problem lacks strong convexity.
- A key technical step involves bounding the difference between dual iterates and the optimal dual solution using the gradient of the dual function and the residual vector.
- The analysis assumes the dual stepsize is sufficiently small to ensure linear convergence, derived via a quadratic bound on the distance to the optimal solution set.
- The method uses a merit function and a mapping M to characterize the relationship between primal-dual iterates and the optimal solution set.
Experimental results
Research questions
- RQ1Does the ADMM converge linearly when applied to problems with more than two blocks of variables?
- RQ2Can linear convergence be established for ADMM without assuming strong convexity in the objective function?
- RQ3What conditions on the dual stepsize ensure linear convergence of ADMM in multi-block and non-strongly convex settings?
- RQ4How can the distance to the optimal solution set be bounded using the proximal residual in the absence of strong convexity?
- RQ5Can error bound techniques be applied to derive global linear convergence for ADMM in general convex separable problems?
Key findings
- The ADMM converges globally linearly for minimizing the sum of any number of convex separable functions, even without strong convexity.
- Linear convergence is established under the condition that the dual stepsize is sufficiently small, which ensures the error bound holds.
- The convergence rate is quantified via a constant τ that depends on problem parameters such as Lipschitz constants and strong convexity parameters of auxiliary functions.
- The result implies linear convergence for LASSO, Group LASSO, and Sparse Group LASSO without requiring strong convexity.
- The analysis provides a theoretical justification for the empirical success of ADMM in multi-block and non-strongly convex problems.
- The proof technique using proximal residuals and error bounds offers a general framework applicable to a wide class of structured convex optimization problems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.