[论文解读] Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry
本文在Kurdyka-Łojasiewicz(KŁ)几何下,首次为非凸-强凹极小化极大优化中的近端梯度下降-上升(GDA)建立了变量收敛性保证。通过引入一种新颖的李雅普诺夫函数,该函数单调递减并引导迭代点收敛至临界点,作者证明了收敛速率取决于KŁ参数,可为线性或次线性,从而解决了非凸极小化极大设置中长期悬而未决的变量收敛性问题。
The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimization, it is important that GDA generates convergent variable sequences rather than convergent sequences of function values or gradient norms. However, the variable convergence of GDA has been proved only under convexity geometries, and there lacks understanding for general nonconvex minimax optimization. This paper fills such a gap by studying the convergence of a more general proximal-GDA for regularized nonconvex-strongly-concave minimax optimization. Specifically, we show that proximal-GDA admits a novel Lyapunov function, which monotonically decreases in the minimax optimization process and drives the variable sequence to a critical point. By leveraging this Lyapunov function and the KŁ geometry that parameterizes the local geometries of general nonconvex functions, we formally establish the variable convergence of proximal-GDA to a critical point $x^*$, i.e., $x_t o x^*, y_t o y^*(x^*)$. Furthermore, over the full spectrum of the KŁ-parameterized geometry, we show that proximal-GDA achieves different types of convergence rates ranging from sublinear convergence up to finite-step convergence, depending on the geometry associated with the KŁ parameter. This is the first theoretical result on the variable convergence for nonconvex minimax optimization.
研究动机与目标
- 解决梯度下降-上升(GDA)在非凸极小化极大优化中是否实现变量收敛的根本性开放问题。
- 刻画由Kurdyka-Łojasiewicz(KŁ)条件参数化的局部函数几何如何影响GDA的收敛速率。
- 为非凸-强凹极小化极大问题中的变量收敛性分析建立新的理论框架。
- 将现有收敛结果从凸-凹或强凸-强凹几何推广至一般非凸设置。
提出的方法
- 为正则化的非凸-强凹极小化极大问题提出一种近端-GDA算法。
- 引入一种新颖的李雅普诺夫函数,使其在整个优化过程中单调递减。
- 利用KŁ几何对局部非凸函数几何进行参数化,并在此通用框架下分析收敛性。
- 通过分析不同KŁ参数区间下李雅普诺夫函数的衰减,推导收敛速率界。
- 利用递推不等式和求和估计,界定迭代点与临界点之间的距离。
- 通过证明当 $ t \to \infty $ 时,$\|x_t - x^*\| \to 0$ 且 $\|y_t - y^*(x^*)\| \to 0$,建立变量收敛性。
实验结果
研究问题
- RQ1GDA在非凸极小化极大优化中是否实现变量收敛,若实现,收敛至何处?
- RQ2由KŁ参数捕获的目标函数局部几何如何影响GDA的收敛速率?
- RQ3能否构造一个李雅普诺夫函数,以确保在非凸-强凹设置中单调递减并引导迭代点收敛至临界点?
- RQ4在不同KŁ参数值下,GDA可实现的收敛速率谱范围是什么?
主要发现
- 在KŁ几何下,近端-GDA算法实现变量收敛至临界点 $x^*$, $y^*(x^*)$,即 $x_t \to x^*$ 且 $y_t \to y^*(x^*)$。
- 收敛速率范围从次线性到有限步收敛,取决于KŁ参数 $\theta$,其中 $\|x_t - x^*\| = \mathcal{O}((t - t_0)^{-\theta/(1 - 2\theta)})$ 对于 $\theta \in (0, \frac{1}{2})$。
- 当 $\theta = \frac{1}{2}$ 时,收敛速率为线性:$\|x_t - x^*\| = \mathcal{O}(\min(2, 1 + \frac{1}{2Mc^2})^{-t/2})$。
- 当KŁ参数满足 $\theta \in (\frac{1}{2}, 1)$ 时,收敛速率提升至有限步收敛,尽管该区间在所提供文本中未详细说明。
- 新颖的李雅普诺夫函数确保了单调递减,并支持在非凸设置中分析变量收敛性。
- 本研究首次为非凸极小化极大优化中的GDA建立了理论上的变量收敛性保证,填补了文献中的关键空白。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。