Skip to main content
QUICK REVIEW

[论文解读] Accelerated Alternating Direction Method of Multipliers: an Optimal $O(1/K)$ Nonergodic Analysis

Huan Li, Zhouchen Lin|arXiv (Cornell University)|Aug 23, 2016
Sparse and Compressive Sensing Techniques参考文献 42被引用 4
一句话总结

本文提出了一种加速的交替方向乘数法(ADMM),其收敛速率最优,达到 O(1/K) 的非ergodic(非平均)收敛率,无需ergodic averaging 即可实现 O(1/K) 的次优性间隙与约束违反度。该方法比ergodic方法更好地保持了稀疏性与低秩结构,是首个针对一般线性约束凸问题实现 O(1/K) 非ergodic ADMM 的方法,在非光滑、非强凸条件下具有最优复杂度。

ABSTRACT

The Alternating Direction Method of Multipliers (ADMM) is widely used for linearly constrained convex problems. It is proven to have an $o(1/\sqrt{K})$ nonergodic convergence rate and a faster $O(1/K)$ ergodic rate after ergodic averaging, which may destroy the sparsity and low-rankness in sparse and low-rank learning, where $K$ is the number of iterations. In this paper, we modify the accelerated ADMM proposed in [Y. Ouyang, Y. Chen, G. Lan, and E. Pasiliao, An Accelerated Linearized Alternating Direction Method of Multipliers, SIAM J. on Imaging Sciences, 2015, 1588-1623] and give an $O(1/K)$ nonergodic convergence rate analysis, which satisfies $|F(x^K)-F(x^*)|\leq O(1/K)$, $\|Ax^K-b\|\leq O(1/K)$ and $x^K$ has a more favorable sparseness and low-rankness than the ergodic result. As far as we know, this is the first $O(1/K)$ nonergodic convergent ADMM type method for general linearly constrained convex problems. Moreover, we show that the lower complexity bound of ADMM type methods for the separable linearly constrained nonsmooth convex problems is $O(1/K)$, which means that our method is optimal.

研究动机与目标

  • 解决传统 ADMM 在线性约束凸问题中非ergodic 收敛率仅为 o(1/√K) 的次优性问题。
  • 克服现有加速 ADMM 方法中因 ergodic averaging 导致的稀疏性与低秩性丧失问题。
  • 开发一种非ergodic 加速 ADMM 变体,实现最优 O(1/K) 收敛率,同时保持有利的解结构。
  • 在非光滑、非强凸条件下,建立 ADMM 类方法的 O(1/K) 复杂度下界,证明所提方法的最优性。

提出的方法

  • 通过引入一种新颖的基于动量的更新策略,对 Ouyang 等人 [18] 提出的加速 ADMM 框架进行改进,实现加速收敛,且无需 ergodic averaging。
  • 采用基于李雅普诺夫函数的分析方法,建立目标函数间隙 |F(x^K) - F(x*)| 与约束违反度 ||Ax^K - b|| 的非ergodic收敛速率。
  • 设计一种参数更新规则,通过精心设计的外推方案平衡原始与对偶步长,确保 O(1/K) 收敛速率。
  • 证明在一般凸、非光滑、非强凸设定下,该方法对目标误差与可行性违反度均实现 O(1/K) 非ergodic 收敛。
  • 通过推导所考虑问题类中 ADMM 类方法的复杂度下界,证明 O(1/K) 速率的最优性。
  • 通过避免平均化操作,保持迭代过程中稀疏性与低秩性,确保最终迭代 x^K 保留机器学习与图像处理中至关重要的结构特性。

实验结果

研究问题

  • RQ1能否使加速 ADMM 在一般线性约束凸问题中实现 O(1/K) 非ergodic 收敛率?
  • RQ2所提方法是否如 ergodic averaging 方法那样,在解迭代中保持稀疏性与低秩性?
  • RQ3在非光滑、非强凸设定下,ADMM 类方法的 O(1/K) 非ergodic 速率是否最优?
  • RQ4收敛分析能否扩展至在实现更快收敛的同时保持稀疏性等结构特性?
  • RQ5对于可分、线性约束、非光滑凸问题,ADMM 类方法的理论复杂度下界是什么?

主要发现

  • 所提出的 ALADMM-NE 在目标函数间隙 |F(x^K) - F(x*)| 与约束违反度 ||Ax^K - b|| 上均实现了 O(1/K) 非ergodic 收敛率。
  • 与 ergodic averaging 不同,该方法在最终迭代 x^K 中保持了稀疏性与低秩性。
  • O(1/K) 非ergodic 收敛率是最优的,本文证明 ADMM 类方法在所考虑问题类中的复杂度下界为 Ω(1/K)。
  • 在组稀疏逻辑回归上的数值实验表明,ALADMM-NE 与 ALADMM-NER 在收敛速度上优于 LADMM 与 ALADMM,同时保持了更优的稀疏性与组稀疏性。
  • 尽管理论界存在限制,该方法在实际中比 ergodic 变体(如 erg-ALADMM)收敛更快,并且在迭代过程中保持了更优的结构特性。
  • 理论分析确认 O(1/K) 速率无法进一步提升,从而确立了所提算法的最优性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。