[Paper Review] Accelerated Alternating Direction Method of Multipliers: an Optimal $O(1/K)$ Nonergodic Analysis
This paper proposes an accelerated Alternating Direction Method of Multipliers (ADMM) with an optimal O(1/K) nonergodic convergence rate, achieving O(1/K) suboptimality gap and constraint violation without ergodic averaging. The method preserves sparsity and low-rank structure better than ergodic counterparts, making it the first O(1/K) nonergodic ADMM for general linearly constrained convex problems and optimal in complexity under nonsmooth, non-strongly convex conditions.
The Alternating Direction Method of Multipliers (ADMM) is widely used for linearly constrained convex problems. It is proven to have an $o(1/\sqrt{K})$ nonergodic convergence rate and a faster $O(1/K)$ ergodic rate after ergodic averaging, which may destroy the sparsity and low-rankness in sparse and low-rank learning, where $K$ is the number of iterations. In this paper, we modify the accelerated ADMM proposed in [Y. Ouyang, Y. Chen, G. Lan, and E. Pasiliao, An Accelerated Linearized Alternating Direction Method of Multipliers, SIAM J. on Imaging Sciences, 2015, 1588-1623] and give an $O(1/K)$ nonergodic convergence rate analysis, which satisfies $|F(x^K)-F(x^*)|\leq O(1/K)$, $\|Ax^K-b\|\leq O(1/K)$ and $x^K$ has a more favorable sparseness and low-rankness than the ergodic result. As far as we know, this is the first $O(1/K)$ nonergodic convergent ADMM type method for general linearly constrained convex problems. Moreover, we show that the lower complexity bound of ADMM type methods for the separable linearly constrained nonsmooth convex problems is $O(1/K)$, which means that our method is optimal.
Motivation & Objective
- Address the suboptimal o(1/√K) nonergodic convergence rate of traditional ADMM for linearly constrained convex problems.
- Overcome the loss of sparsity and low-rankness caused by ergodic averaging in existing accelerated ADMM methods.
- Develop a nonergodic accelerated ADMM variant that achieves the optimal O(1/K) convergence rate while maintaining favorable solution structure.
- Establish the lower complexity bound of O(1/K) for ADMM-type methods in the nonsmooth, non-strongly convex setting, proving optimality of the proposed method.
Proposed method
- Modify the accelerated ADMM framework from Ouyang et al. [18] by introducing a novel momentum-based update strategy to accelerate convergence without ergodic averaging.
- Introduce a Lyapunov function-based analysis to establish nonergodic convergence rates for both objective gap |F(x^K) - F(x*)| and constraint violation ||Ax^K - b||.
- Use a parameter update rule that balances the primal and dual steps via a carefully designed extrapolation scheme, ensuring O(1/K) convergence.
- Prove that the method achieves O(1/K) nonergodic convergence for both objective error and feasibility violation under general convex, nonsmooth, and non-strongly convex settings.
- Demonstrate that the O(1/K) rate is optimal by deriving the lower complexity bound for ADMM-type methods in the considered problem class.
- Preserve sparsity and low-rankness of iterates by avoiding averaging, ensuring the final iterate x^K retains structural properties critical in machine learning and imaging.
Experimental results
Research questions
- RQ1Can an accelerated ADMM achieve an O(1/K) nonergodic convergence rate for general linearly constrained convex problems?
- RQ2Does the proposed method maintain sparsity and low-rankness in the solution iterates, unlike ergodic averaging methods?
- RQ3Is the O(1/K) nonergodic rate optimal for ADMM-type methods in the nonsmooth, non-strongly convex setting?
- RQ4Can the convergence analysis be extended to preserve structural properties like sparsity while achieving faster convergence?
- RQ5What is the theoretical lower bound on the complexity of ADMM-type methods for separable, linearly constrained, nonsmooth convex problems?
Key findings
- The proposed ALADMM-NE achieves an O(1/K) nonergodic convergence rate for both objective gap |F(x^K) - F(x*)| and constraint violation ||Ax^K - b||.
- The method preserves sparsity and low-rankness in the final iterate x^K, unlike ergodic averaging which can destroy such structures.
- The O(1/K) nonergodic rate is optimal, as the paper proves the lower complexity bound for ADMM-type methods is Ω(1/K) under the considered problem class.
- Numerical experiments on group sparse logistic regression show ALADMM-NE and ALADMM-NER outperform LADMM and ALADMM in convergence speed while maintaining superior sparsity and group sparsity.
- The method achieves faster convergence than ergodic variants (e.g., erg-ALADMM) in practice, despite theoretical bounds, and maintains better structural properties in the iterates.
- Theoretical analysis confirms that the O(1/K) rate cannot be improved further, establishing optimality of the proposed algorithm.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.