[论文解读] The limits of min-max optimization algorithms: convergence to spurious non-critical sets
该论文表明,最小-最大优化算法——包括一阶与二阶方法、自适应方案及启发式变体——在非凸/非凹问题中会收敛到虚假的非临界集合,即使这些集合中不包含原博弈的驻点。作者证明,此类算法通常收敛于均值场动力系统的内部链传递(ICT)集合,尽管它们几乎必然避开不稳定流形,但仍可能持续收敛至非临界点的吸引子,揭示了现有算法中固有的收敛失败。
Compared to ordinary function minimization problems, min-max optimization algorithms encounter far greater challenges because of the existence of periodic cycles and similar phenomena. Even though some of these behaviors can be overcome in the convex-concave regime, the general case is considerably more difficult. On that account, we take an in-depth look at a comprehensive class of state-of-the art algorithms and prevalent heuristics in non-convex / non-concave problems, and we establish the following general results: a) generically, the algorithms' limit points are contained in the ICT sets of a common, mean-field system; b) the attractors of this system also attract the algorithms in question with arbitrarily high probability; and c) all algorithms avoid the system's unstable sets with probability 1. On the surface, this provides a highly optimistic outlook for min-max algorithms; however, we show that there exist spurious attractors that do not contain any stationary points of the problem under study. In this regard, our work suggests that existing min-max algorithms may be subject to inescapable convergence failures. We complement our theoretical analysis by illustrating such attractors in simple, two-dimensional, almost bilinear problems.
研究动机与目标
- 分析最先进最小-最大优化算法在非凸/非凹设置下的收敛行为,其中标准收敛保证失效。
- 阐明尽管有理论进展,现有算法为何在简单且近乎双线性的博弈中可能无法收敛至临界点。
- 基于随机逼近理论,统一分析多种算法(SGDA、外梯度法、近端法、自适应方法)的理论框架。
- 证明算法收敛于均值场系统内部链传递(ICT)集合,而这些集合可能不包含临界点,揭示根本性局限。
- 表明即使高级方法如Adam和CGD也可能收敛至劣质解(如最大-最小点),从而削弱其实际可靠性。
提出的方法
- 将最小-最大算法形式化为具有时变步长和噪声梯度的Robbins-Monro(RM)随机逼近方案。
- 通过连续时间均值场动力学中内部链传递(ICT)集合的理论,分析这些算法的渐近行为。
- 为各种算法(如SGDA、SGA、ConO、CGD、OG/PEG)推导其连续时间极限系统,作为离散动力学的均值场近似。
- 利用随机稳定性理论证明,算法以概率1避开不稳定流形,同时收敛至均值场系统的吸引子。
- 应用伪轨迹与扰动理论,将离散算法与连续动力学联系起来,尤其在恒定或自适应步长下。
- 在二维近乎双线性博弈中进行数值实验,以可视化虚假ICT集合,并比较不同方法的算法行为。
实验结果
研究问题
- RQ1在非凸/非凹设置下,最小-最大优化算法是否收敛至原博弈的临界点?
- RQ2现有算法能否收敛至不包含博弈驻点的吸引子?若能,在何种条件下发生?
- RQ3不同算法变体(一阶、二阶、自适应)在收敛至虚假集合方面表现出何种行为?
- RQ4恒定步长与自适应方法(如Adam)在多大程度上加剧了向劣质解的收敛?
- RQ5最小-最大算法的均值场动力学能否用于预测并解释持续循环或收敛至非临界集合的现象?
主要发现
- 最小-最大优化算法通常收敛于同一均值场动力系统的内部链传递(ICT)集合,即使这些集合中不包含原博弈的临界点。
- 所有分析的算法——包括SGDA、SGA、ConO、CGD、OG/PEG以及自适应方法——在简单的二维近乎双线性博弈中均收敛至相同的虚假ICT集合。
- 自适应方法如Adam和ExtraAdam在近乎双线性博弈(11)中收敛至最大-最小点(0,0),该解并非最小-最大均衡,表明其无法找到有意义的解。
- 即使二阶方法如竞争梯度下降(CGD)和共识优化(ConO)也收敛至虚假吸引子,其中ConO唯一收敛至不稳定临界点。
- 这些算法的连续时间均值场动力学在小扰动下保持稳定,表明虚假ICT集合具有鲁棒性,而非数值噪声的产物。
- 尽管几乎必然避开不稳定流形,算法仍被困于非临界点的吸引子中,揭示了非凸最小-最大问题中不可避免的收敛失败。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。