[Paper Review] Linear Convergence of First- and Zeroth-Order Primal-Dual Algorithms for Distributed Nonconvex Optimization
This paper proposes distributed first- and zeroth-order primal-dual algorithms for nonconvex optimization over networks, achieving linear convergence to a global optimum when the global cost function satisfies the Polyak–Łojasiewicz (P–Ł) condition—a weaker requirement than strong convexity. The zeroth-order variant uses a deterministic gradient estimator and matches the convergence rate of the first-order method under identical conditions.
This paper considers the distributed nonconvex optimization problem of minimizing a global cost function formed by a sum of local cost functions by using local information exchange. We first consider a distributed first-order primal-dual algorithm. We show that it converges sublinearly to a stationary point if each local cost function is smooth and linearly to a global optimum under an additional condition that the global cost function satisfies the Polyak-Łojasiewicz condition. This condition is weaker than strong convexity, which is a standard condition for proving linear convergence of distributed optimization algorithms, and the global minimizer is not necessarily unique. Motivated by the situations where the gradients are unavailable, we then propose a distributed zeroth-order algorithm, derived from the considered first-order algorithm by using a deterministic gradient estimator, and show that it has the same convergence properties as the considered first-order algorithm under the same conditions. The theoretical results are illustrated by numerical simulations.
Motivation & Objective
- Address the gap in linear convergence guarantees for distributed nonconvex optimization under weaker conditions than strong convexity.
- Develop a distributed first-order primal-dual algorithm that converges linearly when the global cost function satisfies the P–Ł condition.
- Extend the first-order algorithm to a zeroth-order variant using a deterministic gradient estimator, preserving convergence properties.
- Demonstrate that linear convergence is achievable even when gradients are unavailable, under the same P–Ł condition.
- Validate theoretical results via numerical simulations on a distributed binary classification problem.
Proposed method
- Propose a distributed first-order primal-dual algorithm using local information exchange to minimize a global cost function composed of local functions.
- Establish sublinear convergence to a stationary point for smooth local cost functions, and linear convergence to a global optimum under the P–Ł condition.
- Derive a zeroth-order algorithm by replacing gradients with a deterministic gradient estimator based on function value differences.
- Prove that the zeroth-order algorithm inherits the same convergence properties as the first-order variant under the same assumptions.
- Use Lyapunov functions and contraction arguments to analyze convergence, with key inequalities bounding the distance to the optimal set.
- Validate the theoretical findings through simulations on a nonconvex binary classification problem with 20 agents and 50-dimensional variables.
Experimental results
Research questions
- RQ1Can distributed first-order primal-dual algorithms achieve linear convergence for nonconvex problems under conditions weaker than strong convexity?
- RQ2Does a zeroth-order variant of the first-order algorithm preserve the same convergence rate when gradients are unavailable?
- RQ3How does the P–Ł condition compare to strong convexity in enabling linear convergence for distributed nonconvex optimization?
- RQ4What is the performance trade-off between first-order and zeroth-order methods in terms of function queries and communication rounds?
- RQ5Can deterministic gradient estimation in zeroth-order methods maintain linear convergence under the P–Ł condition?
Key findings
- The proposed first-order primal-dual algorithm converges linearly to a global optimum if the global cost function satisfies the Polyak–Łojasiewicz (P–Ł) condition, which is weaker than strong convexity.
- The zeroth-order algorithm, derived via deterministic gradient estimation, achieves the same linear convergence rate as the first-order method under identical assumptions.
- The convergence rate is linear in the sense that the distance to the optimal set decays exponentially, with the decay rate depending on the P–Ł constant and algorithm parameters.
- Numerical simulations show that the first-order algorithm outperforms state-of-the-art methods like DGD, DFO-GTA, and xFILTER in terms of convergence speed.
- The zeroth-order algorithm exhibits similar early performance to its first-order counterpart but eventually slows to sublinear convergence, highlighting the advantage of gradient access.
- The proposed zeroth-order algorithm achieves better performance than the related DDZO-GTA [65] in terms of function value queries and communication efficiency, as shown in Fig. 2 and Fig. 3.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.