Skip to main content
QUICK REVIEW

[论文解读] Multi-Point Bandit Algorithms for Nonstationary Online Nonconvex Optimization

Abhishek Roy, Krishnakumar Balasubramanian|arXiv (Cornell University)|Jul 31, 2019
Advanced Bandit Algorithms Research参考文献 49被引用 6
一句话总结

本文针对非平稳在线非凸优化问题,提出多点 bandit 算法,引入基于一阶和二阶平稳解及函数值的新型遗憾度量。在低维与高维设置下均建立了次线性遗憾界,包括对弱拟凸函数和弱单调次模函数的常数遗憾结果,并基于二阶高斯 Stein 恒等式提出了一种带 bandit 版本的立方正则化牛顿法以实现 Hessian 矩阵估计。

ABSTRACT

Bandit algorithms have been predominantly analyzed in the convex setting with function-value based stationary regret as the performance measure. In this paper, motivated by online reinforcement learning problems, we propose and analyze bandit algorithms for both general and structured nonconvex problems with nonstationary (or dynamic) regret as the performance measure, in both stochastic and non-stochastic settings. First, for general nonconvex functions, we consider nonstationary versions of first-order and second-order stationary solutions as a regret measure, motivated by similar performance measures for offline nonconvex optimization. In the case of second-order stationary solution based regret, we propose and analyze online and bandit versions of the cubic regularized Newton's method. The bandit version is based on estimating the Hessian matrices in the bandit setting, based on second-order Gaussian Stein's identity. Our nonstationary regret bounds in terms of second-order stationary solutions have interesting consequences for avoiding saddle points in the bandit setting. Next, for weakly quasi convex functions and monotone weakly submodular functions we consider nonstationary regret measures in terms of function-values; such structured classes of nonconvex functions enable one to consider regret measure defined in terms of function values, similar to convex functions. For this case of function-value, and first-order stationary solution based regret measures, we provide regret bounds in both the low- and high-dimensional settings, for some scenarios.

研究动机与目标

  • 为解决非平稳在线非凸优化中缺乏 bandit 算法且缺乏动态遗憾度量的问题。
  • 将平稳遗憾分析从凸设置扩展至非凸场景,尤其针对一般性与结构化非凸函数。
  • 开发适用于 bandit 的方法,通过函数值反馈估计梯度与 Hessian 矩阵,实现在无梯度访问条件下的优化。
  • 提供在低维情况下关于维度呈多项式依赖、在高维稀疏设置下呈多对数依赖的遗憾界。
  • 在随机与非随机反馈下,基于一阶与二阶平稳解及函数值分析非平稳遗憾。

提出的方法

  • 基于梯度大小与二阶平稳解提出非平稳遗憾度量,将离线非凸优化概念扩展至在线 bandit 设置。
  • 提出基于二阶高斯 Stein 恒等式、利用函数值反馈估计 Hessian 矩阵的 bandit 版本立方正则化牛顿法。
  • 通过随机扰动实现多点梯度估计,以在无显式梯度信息时近似梯度与 Hessian 矩阵。
  • 在低维与高维设置下推导遗憾界,后者依赖于结构稀疏性假设以实现关于维度的多对数依赖。
  • 采用带条件期望与鞅差序列的随机逼近框架,控制在线更新中的误差传播。
  • 通过函数光滑性、有界梯度与非平稳性(通过最优点总变差度量)的假设,推导出紧致的遗憾界。

实验结果

研究问题

  • RQ1能否在一般非凸函数的 bandit 设置下,对非平稳遗憾进行有意义的定义与有界化?
  • RQ2在仅具有函数值反馈的 bandit 设置下,如何近似二阶平稳解?
  • RQ3在弱拟凸性或次模性等结构假设下,非平稳非凸 bandit 优化的维度相关遗憾界为何种形式?
  • RQ4bandit 算法能否在非平稳设置下对结构化非凸函数实现常数遗憾?
  • RQ5使用二阶信息(Hessian 估计)如何提升 bandit 非凸优化中的收敛性与鞍点规避性能?

主要发现

  • 本文在梯度大小为基础的遗憾度量下,为非平稳 bandit 优化建立了 O(d√(T + TV_T)) 的遗憾界,该结果在随机与确定性设置下均成立。
  • 对于弱拟凸函数,在稀疏性假设下,本文实现关于维度的多对数遗憾界,从而具备高维适用性。
  • bandit 立方正则化牛顿法在二阶平稳解方面实现了次线性遗憾,且具有鞍点规避的理论保证。
  • 对于单调弱 DR-次模函数,本文基于函数值推导出遗憾界,并显式体现了对曲率参数 γ 的依赖。
  • 在相同假设下,遗憾界在随机与确定性设置间保持不变,表明对反馈噪声具有鲁棒性。
  • 分析表明,通过结构稀疏性假设,维度依赖可从多项式降低至多对数形式,从而实现高维问题的可扩展性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。