Skip to main content
QUICK REVIEW

[論文レビュー] KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints

Aurélien Garivier, Hédi Hadiji|arXiv (Cornell University)|May 14, 2018
Advanced Bandit Algorithms Research参考文献 22被引用数 15
ひとこと要約

本稿では、[0,1] に値をとる確率的マルチアームバンディットにおける、分布依存および分布フリーの両設定で最適なリグレットバウンドを同時に達成する、新しいバンディットアルゴリズム KL-UCB-Switch を提案する。MOSS の分布フリー最適性と KL-UCB の分布依存最適性を組み合わせることで、既知の最もタイトなリグレットバウンドを達成する:最悪分布では $\sqrt{KT}$、問題固有の最適性では $\kappa\ln T$ を達成し、非パラメトリックバンディット理論における長年の未解決問題を解消する。

ABSTRACT

We consider $K$-armed stochastic bandits and consider cumulative regret bounds up to time $T$. We are interested in strategies achieving simultaneously a distribution-free regret bound of optimal order $\sqrt{KT}$ and a distribution-dependent regret that is asymptotically optimal, that is, matching the $κ\ln T$ lower bound by Lai and Robbins (1985) and Burnetas and Katehakis (1996), where $κ$ is the optimal problem-dependent constant. This constant $κ$ depends on the model $\mathcal{D}$ considered (the family of possible distributions over the arms). Ménard and Garivier (2017) provided strategies achieving such a bi-optimality in the parametric case of models given by one-dimensional exponential families, while Lattimore (2016, 2018) did so for the family of (sub)Gaussian distributions with variance less than $1$. We extend this result to the non-parametric case of all distributions over $[0,1]$. We do so by combining the MOSS strategy by Audibert and Bubeck (2009), which enjoys a distribution-free regret bound of optimal order $\sqrt{KT}$, and the KL-UCB strategy by Cappé et al. (2013), for which we provide in passing the first analysis of an optimal distribution-dependent $κ\ln T$ regret bound in the model of all distributions over $[0,1]$. We were able to obtain this non-parametric bi-optimality result while working hard to streamline the proofs (of previously known regret bounds and thus of the new analyses carried out); a second merit of the present contribution is therefore to provide a review of proofs of classical regret bounds for index-based strategies for $K$-armed stochastic bandits.

研究の動機と目的

  • 確率的バンディットにおける分布依存および分布フリーの両レジームで最適なリグレットを達成するバンディットアルゴリズムの設計。
  • 従来、指数型分散族やサブガウス分布などのパラメトリックモデルに限定されてきたバイオプティマル戦略を、[0,1] 上のすべての分布をカバーする非パラメトリックケースに拡張すること。
  • MOSS(最適な分布フリーのリグレット)と KL-UCB(最適な漸近的リグレット)の長所を統合し、1つの適応的戦略として実現すること。
  • インデックスベースのバンディット方策における古典的リグレットバウンドを簡素化・統一し、より明確で統一的な理論的基盤を提供すること。

提案手法

  • 探索と活用のバランスを図るため、信頼区間のしきい値に基づいて MOSS と KL-UCB の間で切り替える戦略を提案する。
  • KL-UCB の部品において、カルバック・ライブーラー(Kullback-Leibler)ダイバージェンスを用いて上側信頼区間を構築し、分布依存レジームにおける漸近的最適性を保証する。
  • MOSS 戦略の信頼区間スケーリングを用いることで、最適な $\sqrt{KT}$ の分布フリーのリグレットバウンドを達成する。
  • [0,1] 値をとるすべての分布の非パラメトリックモデルにおける KL-UCB のリグレットの新たな解析を導入し、この設定で初めて $\kappa\ln T$ の最適な分布依存バウンドを証明する。
  • KL ダイバージェンスの変分表現とラドン=ニコディム微分を用いて、得られた戦略の最適性を証明する。
  • インデックス方策の古典的リグレット証明を包括的かつ簡略化したレビューと再導出を提供し、明確性とアクセスのしやすさを向上させる。

実験結果

リサーチクエスチョン

  • RQ1非パラメトリックモデル([0,1] 上のすべての分布)において、1つのバンディットアルゴリズムが分布依存および分布フリーの両設定で最適なリグレットバウンドを達成できるか。
  • RQ2MOSS(最適なミニマックスリグレット)と KL-UCB(最適な漸近的リグレット)の長所を1つの適応的戦略に統合することは可能か。
  • RQ3[0,1] 上のすべての分布の非パラメトリックモデルにおいて、最悪ケースと問題固有の設定の両方で達成可能な最もタイトなリグレットバウンドは何か。
  • RQ4インデックス方策の古典的リグレット解析を、1つの理論的枠組みの下で簡素化・統一することは可能か。
  • RQ5KL ダイバージェンスと指数型傾き(exponential tilting)は、非パラメトリックバンディットにおける分布依存最適性を達成するために果たす役割は何か。

主な発見

  • KL-UCB-Switch アルゴリズムは、$\sqrt{KT}$ のオーダーの分布フリーのリグレットバウンドを達成し、定数要因を除いて最適である。
  • また、$\kappa\ln T$ のオーダーの分布依存リグレットバウンドも達成し、Lai と Robbins (1985) および Burnetas と Katehakis (1996) が確立した漸近的下界と一致する。
  • このバイオプティマル性は、[0,1] 上のすべての分布の非パラメトリックモデルにおいて、初めて確立されたものであり、従来のパラメトリック族に限定された結果を拡張する。
  • 本稿は、非パラメトリック設定における KL-UCB のリグレットの有限時刻解析を初めて提供し、分布依存レジームにおける最適性を証明する。
  • 理論的解析は簡素化されており、インデックス方策の古典的リグレットバウンドのより明確で統一的な導出を提供する。
  • 変分表現を用いて KL ダイバージェンスの最適性を証明し、ラドン=ニコディム微分とジェンセンの不等式による厳密な正当化を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。