Skip to main content
QUICK REVIEW

[论文解读] Lower complexity bounds of first-order methods for convex-concave bilinear saddle-point problems

Yuyuan Ouyang, Yangyang Xu|arXiv (Cornell University)|Aug 8, 2018
Stochastic Gradient Optimization Techniques参考文献 39被引用 12
一句话总结

本文為求解凸-凹双线性鞍点问题(SPP)的第一类方法建立了紧致的下界复杂度,证明了在一般凸SPP中$O(1/t)$为最优,而在强凸情况下$O(1/t^2)$为最优。作者通过构造最坏情况实例,证明了包括内沃罗夫平滑方案在内的现有加速方法达到了最佳可能的收敛速率,从而在第一类预言机访问下确认了其最优性。

ABSTRACT

On solving a convex-concave bilinear saddle-point problem (SPP), there have been many works studying the complexity results of first-order methods. These results are all about upper complexity bounds, which can determine at most how many efforts would guarantee a solution of desired accuracy. In this paper, we pursue the opposite direction by deriving lower complexity bounds of first-order methods on large-scale SPPs. Our results apply to the methods whose iterates are in the linear span of past first-order information, as well as more general methods that produce their iterates in an arbitrary manner based on first-order information. We first work on the affinely constrained smooth convex optimization that is a special case of SPP. Different from gradient method on unconstrained problems, we show that first-order methods on affinely constrained problems generally cannot be accelerated from the known convergence rate $O(1/t)$ to $O(1/t^2)$, and in addition, $O(1/t)$ is optimal for convex problems. Moreover, we prove that for strongly convex problems, $O(1/t^2)$ is the best possible convergence rate, while it is known that gradient methods can have linear convergence on unconstrained problems. Then we extend these results to general SPPs. It turns out that our lower complexity bounds match with several established upper complexity bounds in the literature, and thus they are tight and indicate the optimality of several existing first-order methods.

研究动机与目标

  • 通过推导大规模SPP的第一类方法的下界复杂度,填补文献中此前仅知上界的关键空白。
  • 确定SPP的第一类方法在收敛速率上是否可进一步改进或已达到最优。
  • 建立与已知上界相匹配的紧致边界,从而确认文献中若干著名算法的最优性。
  • 在统一的信息基础复杂度框架下,分析标准第一类方法与更一般的方法(基于第一类信息生成迭代点)的性能。
  • 研究目标函数信息在可行性与最优性中的作用,特别是在约束优化及不可行性残差边界的情境下。

提出的方法

  • 利用凸二次规划构造最坏情况SPP实例,在过去第一类信息线性张量假设下推导收敛速率的下界。
  • 应用旋转不变性技术,将线性张量方法的下界扩展至仿射约束问题上的一般第一类方法。
  • 运用信息基础复杂度理论,分析为达到给定精度所需的第一类预言机查询的最小数量,将预言机建模为同时返回梯度与矩阵-向量乘积。
  • 证明对于仿射约束的光滑凸问题,$O(1/t)$为最优且无法加速至$O(1/t^2)$,而$O(1/t^2)$为强凸问题的最优速率。
  • 将分析扩展至具有紧致原空间与对偶可行集的一般SPP,推导出与文献中已知上界相匹配的下界。
  • 将推导出的下界与现有上界(如Nesterov 2005,Chen et al. 2014,2017)进行比较,确认对应算法的紧致性与最优性。

实验结果

研究问题

  • RQ1对于凸-凹双线性鞍点问题,第一类方法在一般凸情况下是否可超越$O(1/t)$的收敛速率?
  • RQ2对于强凸SPP,$O(1/t^2)$是否为第一类方法的最佳可能收敛速率,且是否与现有方法的性能一致?
  • RQ3SPP的现有上界复杂度是否代表了根本极限,还是可通过更优的算法设计进一步改进?
  • RQ4下界复杂度如何依赖于问题结构,如目标函数的光滑性及约束矩阵的性质?
  • RQ5能否推导出同时依赖于约束结构与目标函数的下界,特别是针对可行性误差与最优性误差?

主要发现

  • 对于仿射约束的光滑凸问题,$O(1/t)$的收敛速率是最优的,无法加速至$O(1/t^2)$,这与无约束情况相反。
  • 对于强凸问题,$O(1/t^2)$为最佳可能的收敛速率,且该速率是紧致的,证实了内沃罗夫型加速方法的最优性。
  • 对于具有紧致可行集的一般SPP,下界复杂度为凸问题$O(1/t)$,强凸问题$O(1/t^2)$,与Nesterov(2005)及Chen等(2014,2017)的已知上界完全匹配。
  • 定理4.1中的下界与内沃罗夫平滑方案的上界一致,证明了该方法在收敛速率阶上是最优的。
  • 下界结果确认了若干后续方法的最优性,包括Chen等(2014,2017)的方法,其收敛速率阶与推导出的下界完全一致。
  • 结果表明,第一类方法在SPP上的最坏情况复杂度从根本上受限于问题结构与信息获取方式,任何方法都无法超越已建立的边界。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。