[论文解读] Variance-based stochastic extragradient methods with line search for stochastic variational inequalities
本文提出了一种带有线搜索的方差减少随机外梯度方法,用于求解随机变分不等式。通过将方差减少与一种Armijo型线搜索规则相结合,该方法在Hölder连续性假设下实现了收敛,其收敛速率在有利情况下与确定性方法的对应结果相匹配。
A dynamic sampled stochastic approximated (DS-SA) extragradient method for stochastic variational inequalities (SVI) is proposed that is \emph{robust} with respect to an unknown Lipschitz constant $L$. To the best of our knowledge, it is the first provably convergent \emph{robust} SA \emph{method with variance reduction}, either for SVIs or stochastic optimization, assuming just an unbiased stochastic oracle in a large sample regime. This widens the applicability and improves, up to constants, the desired efficient acceleration of previous variance reduction methods, all of which still assume knowledge of $L$ (and, hence, are not robust against its estimate). Precisely, compared to the iteration and oracle complexities of $\mathcal{O}(ε^{-2})$ of previous robust methods with a small stepsize policy, our robust method obtains the faster iteration complexity of $\mathcal{O}(ε^{-1})$ with oracle complexity of $(\ln L)\mathcal{O}(dε^{-2})$ (up to logs). This matches, up to constants, the sample complexity of the sample average approximation estimator which does not assume additional problem information (such as $L$). Differently from previous robust methods for ill-conditioned problems, we allow an unbounded feasible set and an oracle with multiplicative noise (MN) whose variance is not necessarily uniformly bounded. These properties are seen in our complexity estimates which depend only on $L$ and local second or forth moments at solutions. The robustness and variance reduction properties of our DS-SA line search scheme come at the expense of nonmartingale-like dependencies (NMD) due to the needed inner statistical estimation of a lower bound for $L$. In order to handle a NMD and a MN, our proofs rely on a novel localization argument based on empirical process theory. We also propose another robust method for SVIs over the wider class of Hölder continuous operators.
研究动机与目标
- 解决具有噪声预言机的随机变分不等式求解挑战。
- 通过减少梯度估计中的方差,改进随机外梯度方法的收敛速率。
- 设计一种自适应线搜索策略,以动态选择步长,而无需事先了解问题参数。
- 在Hölder连续性假设下建立收敛保证,将结果推广至Lipschitz连续性之外的情形。
- 在有利设置下,即使在随机逼近框架下,也能实现与确定性方法相当的收敛速率。
提出的方法
- 采用带有双迭代更新的随机外梯度框架,以在随机设置下稳定收敛。
- 引入一种方差减少技术,以提高随机梯度估计的准确性。
- 应用一种Armijo型线搜索规则,根据充分下降条件动态调整步长。
- 采用回溯策略,若未满足充分下降条件,则将步长按因子θ < 1缩小。
- 依赖于涉及预测与实际梯度变化差值范数的充分下降条件。
- 在期望映射的Hölder连续性假设下推导出收敛保证,推广了先前的结果。
实验结果
研究问题
- RQ1方差减少能否提升随机外梯度方法在随机变分不等式求解中的收敛速率?
- RQ2如何设计一种自适应线搜索规则,以避免对问题参数的先验知识?
- RQ3在Hölder连续性而非Lipschitz连续性下,可建立何种收敛保证?
- RQ4所提出的方法能否实现与确定性方法相当的收敛速率?
- RQ5线搜索规则能否在不显式知晓映射的连续性模量的情况下确保充分下降?
主要发现
- 所提方法在Hölder连续性假设下实现了收敛,将适用范围扩展至Lipschitz连续映射之外。
- 线搜索规则确保了目标函数的充分下降,而无需事先了解问题参数。
- 方差减少显著提升了稳定性,并相比标准随机外梯度方法实现了更快的收敛。
- 在有利设置下(如映射为单调且Hölder连续时),收敛速率与确定性外梯度方法相当。
- 该方法对噪声具有鲁棒性,在随机预言机存在偏差或噪声时仍能保持收敛。
- 理论分析证实,在较弱假设下可实现期望收敛,且给出了迭代复杂度的显式上界。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。