[論文レビュー] Tsallis-INF: An Optimal Algorithm for Stochastic and Adversarial Bandits
Tsallis-INF は、環境の種別や時間枠に関する事前知識がなくても、確率的および敵対的環境の両方で最適なレジーツバウンドを達成する新しいバンディットアルゴリズムである。オンラインミラー降下と α=1/2 のTsallisエントロピー正則化、および低分散損失推定器を組み合わせることで、確率的および確率的制約付き設定において対数的レジーツを達成し、敵対的レジーツ保証を維持することができる。
We derive an algorithm that achieves the optimal (within constants) pseudo-regret in both adversarial and stochastic multi-armed bandits without prior knowledge of the regime and time horizon. The algorithm is based on online mirror descent (OMD) with Tsallis entropy regularization with power $α=1/2$ and reduced-variance loss estimators. More generally, we define an adversarial regime with a self-bounding constraint, which includes stochastic regime, stochastically constrained adversarial regime (Wei and Luo), and stochastic regime with adversarial corruptions (Lykouris et al.) as special cases, and show that the algorithm achieves logarithmic regret guarantee in this regime and all of its special cases simultaneously with the adversarial regret guarantee.} The algorithm also achieves adversarial and stochastic optimality in the utility-based dueling bandit setting. We provide empirical evaluation of the algorithm demonstrating that it significantly outperforms UCB1 and EXP3 in stochastic environments. We also provide examples of adversarial environments, where UCB1 and Thompson Sampling exhibit almost linear regret, whereas our algorithm suffers only logarithmic regret. To the best of our knowledge, this is the first example demonstrating vulnerability of Thompson Sampling in adversarial environments. Last, but not least, we present a general stochastic analysis and a general adversarial analysis of OMD algorithms with Tsallis entropy regularization for $α\in[0,1]$ and explain the reason why $α=1/2$ works best.
研究の動機と目的
- 環境の種別や時間枠に関する事前知識がなくても、確率的および敵対的環境の両方で最適なレジーツを達成する単一のバンディットアルゴリズムを開発すること。
- 確率的、敵対的、および汚染された確率的バンディット設定の分析を、共通の枠組みで統合すること。
- Tsallisエントロピー正則化の α=1/2 が、多様なバンディット環境において最適なパフォーマンスを実現できることを示すこと。
- Thompson Sampling が敵対的環境に対して脆弱であるのに対し、Tsallis-INF はそうではないことを示すこと。
- Tsallisエントロピーを用いた OMD に対して一般化された分析を行い、α∈[0,1] の範囲で α=1/2 が最適である理由を説明すること。
提案手法
- アルゴリズムは、α=1/2 におけるTsallisエントロピー正則化を用いたオンラインミラー降下(OMD)を用いる。
- 確率的設定におけるレジーツバウンドの向上を図るために、低分散損失推定器を採用する。
- 一般化された確率的および敵対的バンディットを統合する自己バウンド制約の下で手法を分析する。
- 一般化された自己バウンド制約の下で、対数的レジーツを達成し、その特殊ケース(確率的バンディット、確率的制約付き敵対的バンディット、敵対的汚染付き確率的バンディット)に対しても同様に成立する。
- 理論的分析により、α=1/2 が確率的および敵対的両方のレジーツを最小化することが確認される。
- フレームワークは、報酬に基づくデュエルバンディットに対しても拡張され、同様に最適なレジーツを達成する。
実験結果
リサーチクエスチョン
- RQ1単一のバンディットアルゴリズムが、環境の種別に関する事前知識がなくても、確率的および敵対的環境の両方で最適なレジーツを達成できるか?
- RQ2バンディットアルゴリズムにおけるTsallisエントロピー正則化パラメータ α の最適値は何か?
- RQ3UCB1 や Thompson Sampling が失敗する敵対的環境において、Tsallis-INF はどのように動作するか?
- RQ4自己バウンド制約フレームワークは、確率的および敵対的バンディット設定を、1つのレジーツ保証で統合できるか?
- RQ5なぜ OMD における Tsallis エントロピー正則化で α=1/2 が最適なパフォーマンスをもたらすのか?
主な発見
- Tsallis-INF は、環境の種別や時間枠に関する事前知識がなくても、確率的および敵対的バンディット設定の両方で、定数要因の範囲内で最適な(pseudo-regret)を達成する。
- アルゴリズムは、一般化された自己バウンド制約の下で対数的レジーツを達成し、その特殊ケース(確率的バンディット、確率的制約付き敵対的バンディット、敵対的汚染付き確率的バンディット)を含む。
- 実験結果により、Tsallis-INF は確率的環境において UCB1 や EXP3 より顕著に優れていることが示された。
- 敵対的環境では、UCB1 や Thompson Sampling はほぼ線形のレジーツを示すが、Tsallis-INF は対数的レジーツを維持する。
- 本研究は、Thompson Sampling が敵対的環境に対して脆弱であるという、最初の実証的証拠を提供した。
- 理論的分析により、α=1/2 が OMD における Tsallis エントロピー正則化で最適であり、確率的および敵対的両方のレジーツを最小化することが確認された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。