[論文レビュー] First-Order Bayesian Regret Analysis of Thompson Sampling
本稿は、組み合わせ的セミバンドイット問題におけるスケール感受性のある情報比と座標エントロピーを導入することで、Thompson Sampling の情報理論的分析を洗練させ、半バンドイット設定において $\widetilde{O}(\sqrt{dL^*})$ の1次ベイジアンレグレットバウンドを達成した。これは、既存の最良の頻度主義的バウンドと一致する。さらに、閾値付きThompson Samplingを提案し、$L^* \leq \overline{L}^*$ のとき $T$-依存性のないレグレットを達成する。これは、標準的なThompson Samplingが満たさない性質である。
We address online combinatorial optimization when the player has a prior over the adversary's sequence of losses. In this framework, Russo and Van Roy proposed an information-theoretic analysis of Thompson Sampling based on the information ratio, resulting in optimal worst-case regret bounds. In this paper we introduce three novel ideas to this line of work. First we propose a new quantity, the scale-sensitive information ratio, which allows us to obtain more refined first-order regret bounds (i.e., bounds of the form $\sqrt{L^*}$ where $L^*$ is the loss of the best combinatorial action). Second we replace the entropy over combinatorial actions by a coordinate entropy, which allows us to obtain the first optimal worst-case bound for Thompson Sampling in the combinatorial setting. Finally, we introduce a novel link between Bayesian agents and frequentist confidence intervals. Combining these ideas we show that the classical multi-armed bandit first-order regret bound $ ilde{O}(\sqrt{d L^*})$ still holds true in the more challenging and more general semi-bandit scenario. This latter result improves the previous state of the art bound $ ilde{O}(\sqrt{(d+m^3)L^*})$ by Lykouris, Sridharan and Tardos. Moreover we sharpen these results with two technical ingredients. The first leverages a recent insight of Zimmert and Lattimore to replace Shannon entropy with more refined potential functions in the analysis. The second is a \emph{Thresholded} Thompson sampling algorithm, which slightly modifies the original algorithm by never playing low-probability actions. This thresholding results in fully $T$-independent regret bounds when $L^*$ is almost surely upper-bounded, which we show does not hold for ordinary Thompson sampling.
研究の動機と目的
- 組み合わせ的セミバンドイット問題におけるベイジアンと頻度主義的レグレットバウンドのギャップを埋めること。
- 最適損失 $L^*$ に依存するスケーリングを捉えることのできる、Thompson Sampling の洗練された情報理論的分析を構築すること。
- 有界な最適損失 $L^* \leq \overline{L}^*$ の下で $T$-依存性のないレグレットバウンドを確立すること。これは、標準的なThompson Samplingが達成できない性質である。
- ベイジアンエージェントと頻度主義的信頼区間を結びつけることで、ベイジアンと頻度主義的視点を統合すること。
提案手法
- 損失依存のスケーリングを捉える情報比の改良版であるスケール感受性のある情報比を導入すること。
- 組み合わせ的行動全体のシャノンエントロピーを座標エントロピーに置き換えることで、組み合わせ的設定における解析を改善すること。
- シャノンエントロピーを超える洗練された情報理論的解析のため、ツァリスエントロピーと対数バリアントポテンシャル関数を活用すること。
- 低確率の行動を避けることで、有界な $L^*$ におけるロバスト性を確保する、閾値付きThompson Samplingを提案すること。
- 良い腕の事後確率に対する一様な下界を示すために、新しいベイジアン仮説検定の議論を用いること。
- Sionのミニマックス定理を用いて問題をベイジアン設定に還元し、レグレット解析における事前分布の利用を可能にすること。
実験結果
リサーチクエスチョン
- RQ1Thompson Sampling は、半バンドイットの組み合わせ的設定において、形式 $O(\sqrt{dL^*})$ の1次レグレットバウンドを達成できるか?
- RQ2標準的なThompson Samplingアルゴリズムは、$L^* \leq \overline{L}^*$ ほぼ確実に $T$-依存性のないレグレットを達成するか?
- RQ3Thompson Sampling のベイジアン解析を、意味のある方法で頻度主義的信頼区間と結びつけることができるか?
- RQ4情報比フレームワークを、スケール感受性または座標固有の測度を用いて洗練することで、レグレットバウンドを改善できるか?
- RQ5$L^*=0$ のような小損失領域において、$\Omega(d\overline{L}^*)$ のレグレットを回避するThompson Samplingの変種は存在するか?
主な発見
- 本稿は、半バンドイット設定においてThompson Samplingが $\widetilde{O}(\sqrt{dL^*})$ のレグレットを達成することを確立した。これは、既存の最良の頻度主義的バウンドと一致する。
- スケール感受性のある情報比により、損失依存のスケーリングを捉えることで、よりタイトな1次レグレット解析が可能になった。
- 座標エントロピーが、組み合わせ的設定におけるThompson Samplingの最悪ケースレグレットバウンドを初めて最適化する形で、グローバルなシャノンエントロピーに置き換えられた。
- 閾値付きThompson Samplingは、$L^* \leq \overline{L}^*$ のとき $T$-依存性のないレグレットを達成する。これは、標準的なThompson Samplingが満たさない性質である。
- 本稿は、$L^*=0$ であっても $d=O(\sqrt{T})$ であっても、文脈付きバンドイットにおいて標準的なThompson Samplingが高確率で $\Omega(\sqrt{T})$ のレグレットを被る、という事実を証明した。これは、小損失領域で失敗することを示している。
- ベイジアンエージェントと頻度主義的信頼区間との間で、新たなリンクを確立した。これにより、統一された解析フレームワークが可能になった。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。