[论文解读] Complexity Analysis of a Stochastic Cubic Regularisation Method under Inexact Gradient Evaluations and Dynamic Hessian Accuracy.
本文提出了一种带有自适应、不精确黑塞矩阵和梯度近似值的随机立方正则化方法,用于非凸优化。该方法将确定性复杂度分析扩展至随机设置,证明了在 O(ε^(-3/2)) 期望迭代次数内收敛至一阶平稳点,该结果与在不精确性及有界误差方差条件下已知的最坏情况最优复杂度相匹配。
We here adapt an extended version of the adaptive cubic regularisation method with dynamic inexact Hessian information for nonconvex optimisation in [2] to the stochastic optimisation setting. While exact function evaluations are still considered, this novel variant inherits the innovative use of adaptive accuracy requirements for Hessian approximations introduced in [2] and additionally employs inexact computations of the gradient. Without restrictions on the variance of the errors, we assume that these approximations are available within a sufficiently large, but fixed, probability and we extend, in the spirit of [13], the deterministic analysis of the framework to its stochastic counterpart, showing that the expected number of iterations to reach a first-order stationary point matches the well known worst-case optimal complexity. This is, in fact, still given by O(epsilon^(-3/2)), with respect to the first-order epsilon tolerance.
研究动机与目标
- 将带有动态黑塞矩阵精度的自适应立方正则化框架扩展至随机优化设置。
- 解决在非凸优化中不精确梯度评估的挑战,且不施加误差方差限制。
- 在不精确性和概率近似保证下,保持最优最坏情况迭代复杂度。
- 将 [2] 的确定性复杂度分析推广至带有不精确梯度和黑塞矩阵近似的随机设置。
- 证明达到一阶平稳点的期望迭代次数与已知最优复杂度界一致。
提出的方法
- 通过引入不精确的梯度评估,将 [2] 中的自适应立方正则化方法扩展至随机优化。
- 引入对黑塞矩阵近似的动态精度要求,根据向最优解的进展程度动态调整精度。
- 假设梯度近似值以固定且足够大的概率可用,且不施加方差限制。
- 将 [13] 中的确定性复杂度框架扩展至随机情形,同时保持理论保证。
- 利用近似误差的概率界,确保在不精确性下的收敛性。
- 采用类似信赖域的机制结合立方正则化,以控制步长并确保目标函数的充分下降。
实验结果
研究问题
- RQ1是否可以将带有动态黑塞矩阵精度的自适应立方正则化方法扩展至具有不精确梯度的随机设置?
- RQ2在不精确梯度和黑塞矩阵近似下,所提随机方法的期望迭代复杂度是多少?
- RQ3尽管存在不精确性,该方法是否仍能保持最优的 O(ε^(-3/2)) 复杂度界?
- RQ4近似质量的概率保证如何影响收敛性和复杂度?
- RQ5能否严格地将带有不精确梯度的确定性复杂度分析扩展至随机情形?
主要发现
- 所提出的随机立方正则化方法在达到一阶平稳点时,实现了 O(ε^(-3/2)) 的期望迭代复杂度。
- 即使在不精确的梯度和黑塞矩阵评估下,该复杂度界仍与非凸优化中已知的最坏情况最优复杂度一致。
- 该方法保持了对黑塞矩阵近似的自适应精度要求,其精度根据进展动态调整。
- 该分析无需对梯度近似误差方差施加限制,仅依赖于近似值以固定且足够大的概率可用这一条件。
- 该框架成功地将 [2] 和 [13] 的确定性复杂度分析推广至带有不精确性的随机设置。
- 期望迭代次数相对于一阶容差 ε 的尺度是最优的,证实了在不精确性下的鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。