[论文解读] Linear Convergence of First- and Zeroth-Order Primal-Dual Algorithms for Distributed Nonconvex Optimization
本文提出了一种用于网络中非凸优化的分布式一阶和零阶原始-对偶算法,在全局代价函数满足Polyak–Łojasiewicz(P–Ł)条件时,实现了线性收敛至全局最优解——该条件弱于强凸性。零阶变体使用确定性梯度估计器,其收敛速率与相同条件下的一阶方法一致。
This paper considers the distributed nonconvex optimization problem of minimizing a global cost function formed by a sum of local cost functions by using local information exchange. We first consider a distributed first-order primal-dual algorithm. We show that it converges sublinearly to a stationary point if each local cost function is smooth and linearly to a global optimum under an additional condition that the global cost function satisfies the Polyak-Łojasiewicz condition. This condition is weaker than strong convexity, which is a standard condition for proving linear convergence of distributed optimization algorithms, and the global minimizer is not necessarily unique. Motivated by the situations where the gradients are unavailable, we then propose a distributed zeroth-order algorithm, derived from the considered first-order algorithm by using a deterministic gradient estimator, and show that it has the same convergence properties as the considered first-order algorithm under the same conditions. The theoretical results are illustrated by numerical simulations.
研究动机与目标
- 填补在弱于强凸性条件下的分布式非凸优化线性收敛性保证的空白。
- 设计一种分布式一阶原始-对偶算法,在全局代价函数满足P–Ł条件时实现线性收敛。
- 将一阶算法扩展为使用确定性梯度估计器的零阶变体,同时保持收敛特性。
- 证明即使在梯度不可用的情况下,只要满足相同的P–Ł条件,线性收敛依然可实现。
- 通过在分布式二分类问题上的数值仿真验证理论结果。
提出的方法
- 提出一种基于本地信息交换的分布式一阶原始-对偶算法,用于最小化由局部函数组成的全局代价函数。
- 为光滑的局部代价函数建立次线性收敛至驻点的结果,并在P–Ł条件下实现向全局最优的线性收敛。
- 通过用基于函数值差值的确定性梯度估计器替代梯度,推导出零阶算法。
- 证明在相同假设下,零阶算法继承了一阶变体的相同收敛特性。
- 使用李雅普诺夫函数和压缩论证分析收敛性,关键不等式用于界定到最优解集的距离。
- 通过在20个代理和50维变量的非凸二分类问题上的仿真验证理论发现。
实验结果
研究问题
- RQ1在弱于强凸性的条件下,分布式一阶原始-对偶算法能否实现非凸问题的线性收敛?
- RQ2当梯度不可用时,一阶算法的零阶变体是否能保持相同的收敛速率?
- RQ3P–Ł条件与强凸性相比,在实现分布式非凸优化线性收敛方面有何差异?
- RQ4在函数查询次数和通信轮次方面,一阶与零阶方法之间的性能权衡如何?
- RQ5在P–Ł条件下,零阶方法中的确定性梯度估计是否能维持线性收敛?
主要发现
- 所提出的**一阶原始-对偶算法**在全局代价函数满足Polyak–Łojasiewicz(P–Ł)条件时,可实现向全局最优的线性收敛,该条件弱于强凸性。
- 通过确定性梯度估计导出的**零阶算法**在相同假设下,实现了与一阶方法相同的线性收敛速率。
- 收敛速率是线性的,即到最优解集的距离呈指数衰减,衰减速率取决于P–Ł常数和算法参数。
- 数值仿真表明,一阶算法在收敛速度方面优于DGD、DFO-GTA和xFILTER等先进方法。
- 零阶算法在初期性能与一阶方法相似,但最终收敛速度减缓至次线性,凸显了梯度信息的优势。
- 所提出的零阶算法在函数值查询次数和通信效率方面优于相关方法DDZO-GTA [65],如图2和图3所示。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。