[论文解读] On the Sublinear Convergence Rate of Multi-Block ADMM
该论文在较弱条件下首次建立了标准多块 ADMM(N ≥ 3)的次线性收敛速率:其中一个函数为凸函数,其余 N−1 个函数为强凸函数,且惩罚参数位于特定范围内。论文证明了 O(1/t) 的递推平均收敛速率和 o(1/t) 的非递推平均收敛速率,解决了实际多块 ADMM 应用中长期存在的收敛性不确定性问题。
The alternating direction method of multipliers (ADMM) is widely used in solving structured convex optimization problems. Despite of its success in practice, the convergence properties of the standard ADMM for minimizing the sum of $N$ $(N\geq 3)$ convex functions with $N$ block variables linked by linear constraints, have remained unclear for a very long time. In this paper, we present convergence and convergence rate results for the standard ADMM applied to solve $N$-block $(N\geq 3)$ convex minimization problem, under the condition that one of these functions is convex (not necessarily strongly convex) and the other $N-1$ functions are strongly convex. Specifically, in that case the ADMM is proven to converge with rate $O(1/t)$ in a certain ergodic sense, and $o(1/t)$ in non-ergodic sense, where $t$ denotes the number of iterations. As a by-product, we also provide a simple proof for the $O(1/t)$ convergence rate of two-block ADMM in terms of both objective error and constraint violation, without assuming any condition on the penalty parameter and strong convexity on the functions.
研究动机与目标
- 为标准多块 ADMM(N ≥ 3)的长期悬而未决的收敛性问题提供解决方案,该方法在实践中表现成功但理论分析尚不明确。
- 识别出标准 ADMM 在 N ≥ 3 时收敛且具有可量化次线性速率的充分条件。
- 为两块 ADMM 的 O(1/t) 收敛速率提供一种简单统一的证明,无需强凸性假设或惩罚参数约束。
- 在递推和非递推两种意义下,为多块 ADMM 建立收敛性保证,填补了理论理解中的关键空白。
提出的方法
- 提出一种新颖的李雅普诺夫函数,基于原始可行性、对偶不一致性与目标函数差距的加权和,专为多块 ADMM 设计。
- 引入一个关键不等式(引理 3.1),利用子问题的最优性条件,界定了每轮迭代中李雅普诺夫函数的下降量。
- 推导出李雅普诺夫函数的递推不等式,形成可伸缩求和形式,从而实现收敛速率分析。
- 通过递推迭代平均实现目标函数误差和约束违反度的 O(1/t) 收敛速率。
- 应用非递推分析方法,证明了 o(1/t) 的收敛速率,该速率在某些情况下优于递推速率。
- 通过惩罚参数和强凸性参数界定了对偶间隙与可行性违反度,从而确立了收敛速率。
实验结果
研究问题
- RQ1在何种条件下,可保证标准多块 ADMM(N ≥ 3)以次线性速率收敛?
- RQ2是否可以在不假设所有函数均为强凸函数的前提下分析多块 ADMM 的收敛速率?
- RQ3惩罚参数与多块 ADMM 收敛性之间的确切关系是什么?
- RQ4能否为两块 ADMM 的 O(1/t) 收敛速率提供一种无需强凸性假设的简洁证明?
- RQ5在多块 ADMM 设置下,递推与非递推收敛速率如何比较?
主要发现
- 当一个函数为凸函数,其余 N−1 个函数为强凸函数,且惩罚参数位于特定有界区域内时,标准多块 ADMM 以 O(1/t) 的递推平均速率收敛。
- 非递推迭代点的收敛速率为 o(1/t),在逐点收敛意义下快于递推速率。
- 通过精心构造的李雅普诺夫函数,同时捕捉原始可行性、对偶不一致性与目标函数差距,建立了收敛速率。
- 为两块 ADMM 的 O(1/t) 收敛速率提供了一种简洁证明,无需强凸性假设或惩罚参数约束。
- 分析结果表明,标准 ADMM 在较弱条件下可被理论保证收敛,解释了其在实践中的成功表现。
- 研究结果为多块 ADMM 在实践中被广泛使用提供了理论基础,尽管此前存在其收敛性不成立的反例。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。