Skip to main content
QUICK REVIEW

[论文解读] Global Convergence of Arbitrary-Block Gradient Methods for Generalized Polyak-Łojasiewicz Functions

Dominik Csiba, Peter Richtárik|arXiv (Cornell University)|Sep 9, 2017
Stochastic Gradient Optimization Techniques参考文献 11被引用 15
一句话总结

本文提出了一套统一框架,用于分析非凸优化问题中任意块梯度方法的性能,其核心创新在于引入了块选择规则的占比函数(proportion function)与弱Polyak-Łojasiewicz(WPL)条件。该框架在光滑与非光滑情形下,统一了多种非凸函数的全局收敛性分析,并为贪婪小批量选择等新型策略提供了新颖的收敛速率,实现了理论上的统一。

ABSTRACT

In this paper we introduce two novel generalizations of the theory for gradient descent type methods in the proximal setting. First, we introduce the proportion function, which we further use to analyze all known (and many new) block-selection rules for block coordinate descent methods under a single framework. This framework includes randomized methods with uniform, non-uniform or even adaptive sampling strategies, as well as deterministic methods with batch, greedy or cyclic selection rules. Second, the theory of strongly-convex optimization was recently generalized to a specific class of non-convex functions satisfying the so-called Polyak-Łojasiewicz condition. To mirror this generalization in the weakly convex case, we introduce the Weak Polyak-Łojasiewicz condition, using which we give global convergence guarantees for a class of non-convex functions previously not considered in theory. Additionally, we establish (necessarily somewhat weaker) convergence guarantees for an even larger class of non-convex functions satisfying a certain smoothness assumption only. By combining the two abovementioned generalizations we recover the state-of-the-art convergence guarantees for a large class of previously known methods and setups as special cases of our general framework. Moreover, our frameworks allows for the derivation of new guarantees for many new combinations of methods and setups, as well as a large class of novel non-convex objectives. The flexibility of our approach offers a lot of potential for future research, as a new block selection procedure will have a convergence guarantee for all objectives considered in our framework, while a new objective analyzed under our approach will have a whole fleet of block selection rules with convergence guarantees readily available.

研究动机与目标

  • 将梯度型方法的收敛性理论从强凸性和PL条件推广至弱凸与非凸设置。
  • 利用占比函数,将随机化、确定性、自适应等多种块选择策略统一于同一分析框架之下。
  • 为满足弱Polyak-Łojasiewicz(WPL)条件的一类新型非凸函数,建立全局收敛性保证。
  • 在WPL条件与光滑性假设下,推导出新型块选择规则(如贪婪小批量)的收敛速率。
  • 实现收敛性保证的可迁移性:新块选择规则可继承框架中所有目标的收敛性保证,反之亦然。

提出的方法

  • 引入占比函数,以统一方式分析所有已知及新型块选择规则(均匀、非均匀、自适应、循环、贪婪等)。
  • 提出弱Polyak-Łojasiewicz(WPL)条件,作为经典PL不等式在弱凸函数上的推广,其定义为 $\|\nabla f(\mathbf{x})\| \cdot \|\mathbf{x}-\mathbf{x}^*\| \geq \sqrt{\mu} \cdot \xi(\mathbf{x})$。
  • 基于占比函数与强迫函数,建立通用下降引理,用于分析在不同块选择规则下的收敛性。
  • 在WPL条件下,推导出光滑与非光滑问题的收敛速率,包括贪婪小批量情形下的 $K \geq \frac{\xi(\mathbf{x}^0)}{\epsilon \lambda_{\min}(\mathbb{E}[\mathbf{M}_{[S]}^{-1}])} \log\left(\frac{\xi(\mathbf{x}^0)}{\epsilon}\right)$。
  • 将该框架应用于已知方法,恢复其最先进的收敛速率,并为此前未被分析的组合(方法与目标)推导出新的保证。
  • 通过数值实验验证了在光滑与非光滑设置下的全局收敛性,以及在WPL条件下梯度模长的局部收敛性。

实验结果

研究问题

  • RQ1能否通过单一分析框架,统一任意块选择规则下的块坐标下降收敛性理论?
  • RQ2弱Polyak-Łojasiewicz条件是否能将全局收敛性保证扩展至超越经典PL条件的更广泛非凸函数类?
  • RQ3能否对新型块选择策略(如贪婪小批量)进行严格分析,并证明其可实现具有竞争力的收敛速率?
  • RQ4在WPL条件下,光滑与非光滑问题的收敛速率如何?与现有方法相比有何差异?
  • RQ5该框架是否能在无需从头推导证明的前提下,为新组合的块选择规则与非凸目标提供收敛性保证?

主要发现

  • 弱Polyak-Łojasiewicz(WPL)条件将经典PL不等式推广至弱凸函数,使更广泛的非凸问题实现全局收敛。
  • 在WPL条件下,光滑问题中贪婪小批量选择的收敛速率为 $K \geq \frac{\xi(\mathbf{x}^0)}{\epsilon \lambda_{\min}(\mathbb{E}[\mathbf{M}_{[S]}^{-1}])} \log\left(\frac{\xi(\mathbf{x}^0)}{\epsilon}\right)$,这是该采样策略的全新结果。
  • 对于非光滑问题,占比函数导出的收敛速率为 $K \geq \frac{\xi(\mathbf{x}^0)nL_{\tau}}{\tau\epsilon} \log\left(\frac{\xi(\mathbf{x}^0)}{\epsilon}\right)$,在WPL框架下提供了新的收敛保证。
  • 该框架可作为特例恢复已知方法(如随机、循环、贪婪)的最先进收敛速率,体现了其通用性。
  • 数值实验验证了串行坐标下降在光滑与非光滑非凸问题上的全局收敛性,以及在WPL条件下梯度模长的局部收敛性。
  • 理论保证在 $K$ 次迭代内,要么达到最优性间隙 $\xi(\mathbf{x}^K) \leq \epsilon$,要么达到梯度范数 $\lambda(\mathbf{x}^k) \leq \epsilon$,确保了在非凸场景下的实际收敛性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。