Skip to main content
QUICK REVIEW

[论文解读] KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints

Aurélien Garivier, Hédi Hadiji|arXiv (Cornell University)|May 14, 2018
Advanced Bandit Algorithms Research参考文献 22被引用 15
一句话总结

本文提出KL-UCB-Switch,一种新型的bandit算法,在奖励取值于[0,1]的随机多臂bandit问题中,同时在分布依赖和分布无关两种设定下实现了最优后悔界。通过结合MOSS的分布无关最优性与KL-UCB的分布依赖最优性,该算法达到了目前已知最紧的后悔界:最坏情况下的$ olimits textrm{regret} = \sqrt{KT}$,以及问题特定最优性下的$ olimits textrm{regret} = \kappa\ln T$,从而解决了非参数bandit理论中长期存在的开放问题。

ABSTRACT

We consider $K$-armed stochastic bandits and consider cumulative regret bounds up to time $T$. We are interested in strategies achieving simultaneously a distribution-free regret bound of optimal order $\sqrt{KT}$ and a distribution-dependent regret that is asymptotically optimal, that is, matching the $κ\ln T$ lower bound by Lai and Robbins (1985) and Burnetas and Katehakis (1996), where $κ$ is the optimal problem-dependent constant. This constant $κ$ depends on the model $\mathcal{D}$ considered (the family of possible distributions over the arms). Ménard and Garivier (2017) provided strategies achieving such a bi-optimality in the parametric case of models given by one-dimensional exponential families, while Lattimore (2016, 2018) did so for the family of (sub)Gaussian distributions with variance less than $1$. We extend this result to the non-parametric case of all distributions over $[0,1]$. We do so by combining the MOSS strategy by Audibert and Bubeck (2009), which enjoys a distribution-free regret bound of optimal order $\sqrt{KT}$, and the KL-UCB strategy by Cappé et al. (2013), for which we provide in passing the first analysis of an optimal distribution-dependent $κ\ln T$ regret bound in the model of all distributions over $[0,1]$. We were able to obtain this non-parametric bi-optimality result while working hard to streamline the proofs (of previously known regret bounds and thus of the new analyses carried out); a second merit of the present contribution is therefore to provide a review of proofs of classical regret bounds for index-based strategies for $K$-armed stochastic bandits.

研究动机与目标

  • 设计一种bandit算法,使其在随机bandit问题的分布依赖与分布无关两种情形下均实现最优后悔。
  • 将先前仅限于指数族或次高斯分布等参数模型的双最优策略,扩展至[0,1]上所有分布的非参数情形。
  • 将MOSS(最优分布无关后悔)与KL-UCB(最优渐近后悔)的优势统一为单一自适应策略。
  • 简化并重新推导基于索引的bandit策略的经典后悔界,提供更清晰、更统一的理论基础。

提出的方法

  • 基于置信度阈值,提出在MOSS与KL-UCB之间切换的策略,以平衡探索与利用。
  • 在KL-UCB部分利用Kullback-Leibler散度构造置信上界,确保在分布依赖设定下的渐近最优性。
  • 采用MOSS策略的置信区间缩放方式,实现最优的$\sqrt{KT}$分布无关后悔界。
  • 提出对KL-UCB在所有[0,1]取值分布的非参数模型下的后悔界的新分析,首次证明了该设定下最优的$\kappa\ln T$分布依赖后悔界。
  • 应用KL散度的变分表示与Radon-Nikodym导数,证明所提策略的最优性。
  • 提供经典索引策略后悔界证明的自包含综述与简化,提升清晰度与可及性。

实验结果

研究问题

  • RQ1是否存在一种单一bandit算法,能在[0,1]上所有分布的非参数模型中,同时在分布依赖与分布无关设定下实现最优后悔界?
  • RQ2是否可能将MOSS(最优最小最大后悔)与KL-UCB(最优渐近后悔)的优势整合为单一自适应策略?
  • RQ3在[0,1]上所有分布的非参数模型中,最坏情况与问题特定设定下可实现的最紧后悔界是什么?
  • RQ4经典索引策略的后悔分析能否在单一理论框架下被简化并统一?
  • RQ5KL散度与指数倾斜在非参数bandit中实现分布依赖最优性中起什么作用?

主要发现

  • KL-UCB-Switch算法实现了$\sqrt{KT}$量级的分布无关后悔界,该界在常数因子意义下最优。
  • 同时,其分布依赖后悔界达到$\kappa\ln T$量级,与Lai和Robbins(1985)以及Burnetas和Katehakis(1996)建立的渐近下界完全匹配。
  • 这是首次在[0,1]上所有分布的非参数模型中实现双最优性,扩展了以往仅限于参数族的研究结果。
  • 本文首次提供了KL-UCB在非参数设定下的有限时间后悔分析,证明其在分布依赖设定下的最优性。
  • 理论分析得到简化,提供了对索引策略经典后悔界更清晰、更统一的推导。
  • 通过KL散度的变分表示,结合Radon-Nikodym导数与Jensen不等式,严谨证明了所提策略的最优性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。