Skip to main content
QUICK REVIEW

[论文解读] A Proximal Alternating Direction Method of Multiplier for Linearly Constrained Nonconvex Minimization

Jiawei Zhang, Zhi‐Quan Luo|arXiv (Cornell University)|Dec 26, 2018
Sparse and Compressive Sensing Techniques参考文献 24被引用 7
一句话总结

本文提出了一种用于线性约束非凸优化的近端交替方向乘子法(ADMM),通过引入平滑化的原始变量迭代序列和增广拉格朗日函数中的二次近端项来稳定振荡行为。该方法建立了全局收敛至驻点的结果,迭代复杂度提升至 𝒪(1/ε²),优于先前最优的 𝒪(1/ε³),并利用新型势函数与误差界分析,证明了在二次目标函数下具有线性收敛性。

ABSTRACT

Consider the minimization of a nonconvex differentiable function over a polyhedron. A popular primal-dual first-order method for this problem is to perform a gradient projection iteration for the augmented Lagrangian function and then update the dual multiplier vector using the constraint residual. However, numerical examples show that this approach can exhibit "oscillation" and may not converge. In this paper, we propose a proximal alternating direction method of multipliers for the multi-block version of this problem. A distinctive feature of this method is the introduction of a "smoothed" (i.e., exponentially weighted) sequence of primal iterates, and the inclusion, at each iteration, to the augmented Lagrangian function a quadratic proximal term centered at the current smoothed primal iterate. The resulting proximal augmented Lagrangian function is inexactly minimized (via a gradient projection step) at each iteration while the dual multiplier vector is updated using the residual of the linear constraints. When the primal and dual stepsizes are chosen sufficiently small, we show that suitable "smoothing" can stabilize the "oscillation", and the iterates of the new proximal ADMM algorithm converge to a stationary point under some mild regularity conditions. Furthermore, when the objective function is quadratic, we establish the linear convergence of the algorithm. Our proof is based on a new potential function and a novel use of error bounds.

研究动机与目标

  • 解决标准ADMM在非凸、线性约束问题中出现的不稳定与收敛性不足问题。
  • 通过引入平滑化的原始变量迭代序列,稳定原始-对偶梯度方法中常见的振荡行为。
  • 开发一种一阶方法,实现寻找ε-驻点解的改进迭代复杂度。
  • 在较弱的正则性条件下,建立全局收敛至驻点的理论保证。
  • 利用新型势函数与误差界分析技术,证明在目标函数为二次函数时具有线性收敛性。

提出的方法

  • 引入一种指数加权的平滑原始变量迭代序列,以增强算法稳定性。
  • 在每次迭代中,将中心位于当前平滑迭代点的二次近端项加入增广拉格朗日函数。
  • 通过梯度投影步骤对近端增广拉格朗日函数进行不精确最小化。
  • 利用线性约束的残差更新对偶乘子。
  • 采用新型势函数与误差界分析技术,证明收敛性。
  • 采用精心设计的步长策略,确保原始变量与对偶变量步长足够小,以保障稳定性与收敛性。

实验结果

研究问题

  • RQ1是否可以通过近端ADMM变体稳定非凸、线性约束优化中的振荡行为?
  • RQ2对于非凸问题,寻找ε-驻点解的近端ADMM的迭代复杂度是多少?
  • RQ3在较弱正则性条件下,所提方法是否能实现全局收敛至驻点?
  • RQ4当目标函数为二次函数时,能否建立线性收敛性?
  • RQ5在实际性能上,该新方法相较于标准ADMM在收敛速度与效率方面表现如何?

主要发现

  • 当原始变量与对偶变量步长足够小时,所提近端ADMM在较弱正则性条件下可实现全局收敛至驻点。
  • 寻找ε-驻点解的迭代复杂度为 𝒪(1/ε²),优于该类问题中已知最优复杂度 𝒪(1/ε³)。
  • 对于二次目标函数,算法表现出线性收敛性,提供了强有力的全局收敛保证。
  • 数值实验表明,所提方法显著优于原始的双重循环ADMM,所需梯度评估次数大幅减少。
  • 收敛性分析依赖于新型势函数与精细化的误差界技术,从而获得更强的理论保证。
  • 通过引入平滑原始变量迭代与近端正则化,该方法有效稳定了标准ADMM中观察到的振荡行为。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。