Skip to main content
QUICK REVIEW

[論文レビュー] Diffusion Approximations for Thompson Sampling in the Small Gap Regime

Lin Fan, Peter W. Glynn|arXiv (Cornell University)|May 19, 2021
Advanced Bandit Algorithms Research参考文献 53被引用数 7
ひとこと要約

本稿は、腕の平均差が $1/\sqrt{n}$ のスケールで変化する小差領域におけるThompsonサンプリングの弱収束理論を構築する。アルゴリズムのダイナミクスが確率微分方程式(SDE)および確率的常微分方程式(random ODE)の解に弱収束することを示し、主な貢献は、報酬分布や事前分布に依存しない普遍的な拡散極限を示す不変性原理の確立である。この極限は、正規分布に対するものと同一である。

ABSTRACT

We study the process-level dynamics of Thompson sampling in the ``small gap'' regime. The small gap regime is one in which the gaps between the arm means are of order $\sqrtγ$ or smaller and the time horizon is of order $1/γ$, where $γ$ is small. As $γ\downarrow 0$, we show that the process-level dynamics of Thompson sampling converge weakly to the solutions to certain stochastic differential equations and stochastic ordinary differential equations. Our weak convergence theory is developed from first principles using the Continuous Mapping Theorem, can handle stationary, weakly dependent reward processes, and can also be adapted to analyze a variety of sampling-based bandit algorithms. Indeed, we show that the process-level dynamics of many sampling-based bandit algorithms -- including Thompson sampling designed for any single-parameter exponential family of rewards, as well as non-parametric bandit algorithms based on bootstrap re-sampling -- satisfy an invariance principle. Namely, their weak limits coincide with that of Gaussian parametric Thompson sampling with Gaussian priors. Moreover, in the small gap regime, the regret performance of these algorithms is generally insensitive to model mis-specification, changing continuously with increasing degrees of mis-specification.

研究の動機と目的

  • 小差領域におけるThompsonサンプリングの解析のための厳密な弱収束フレームワークを構築すること。
  • 腕の平均差が $\Delta/\sqrt{n}$ にスケーリングされる条件下で、時間枠 $n \to \infty$ におけるThompsonサンプリングダイナミクスの漸近的挙動を特定すること。
  • $1/\sqrt{n}$ スケーリング下で、異なる報酬分布や事前分布に対しても同じ極限に収束する普遍的な拡散過程が得られることを確立すること。
  • 弱収束理論を、有限または無限の行動集合、時間的に変化する行動集合を含む、マルチアームバンディットおよび線形バンディット設定へと拡張すること。
  • 連続写像定理を用いた第一原理的導出を提供し、他のサンプリングベースのバンディットアルゴリズムへの適応を可能にすること。

提案手法

  • 腕の平均差の $\Delta/\sqrt{n}$ スケーリング下で、Thompsonサンプリングのダイナミクスを弱収束(分布収束)の観点から分析する。
  • 有限な $n$ におけるThompsonサンプリング行動を記述する離散時間版のSDEおよび確率的ODEを導出する。
  • $D_{\mathbb{R}^d}[0,1]$ における連続写像定理と緊張性基準を用いて、連続な拡散極限への弱収束を証明する。
  • 経験プロセスの収束を検証するため、Glivenko-Cantelli定理とブレケット化エントロピー条件を用いる。
  • マルティングール関数中心極限定理を用いて、指定された共分散構造を持つブラウン運動への収束を確立する。
  • $1/\sqrt{n}$ スケーリング下で、$D_{\mathbb{R}}[0,1]$ における緊張性の十分条件を、モーメントバウンドと増分分散の制御により得る。

実験結果

リサーチクエスチョン

  • RQ1腕の平均差が $1/\sqrt{n}$ にスケーリングされるとき、Thompsonサンプリングのダイナミクスは漸近的にどのように振る舞うか?
  • RQ2離散的時間のThompsonサンプリングダイナミクスは、分布収束の意味でどの連続時間確率過程に収束するか?
  • RQ3Thompsonサンプリングの弱拡散極限は、異なる報酬分布や事前分布に対して普遍的か?
  • RQ4弱収束フレームワークは、有限または無限の行動集合、時間的に変化する行動集合を含む線形バンディットに拡張可能か?
  • RQ5事後分布の近似とブートストラップサンプリングは、極限の拡散過程にどのように寄与するか?

主な発見

  • $\Delta/\sqrt{n}$ スケーリング下で、Thompsonサンプリングのダイナミクスは $n \to \infty$ の下でSDEおよび確率的ODEの解に弱収束する。
  • 極限の拡散過程は普遍的である:特定の報酬分布や事前分布に依存せず、事後分布が適切に近似されれば同じ極限に収束する。
  • 弱拡散極限は、正規分布の報酬と事前分布に対するものと一致し、古典的なBernstein-von Mises近似を裏付ける。
  • 不変性原理が成り立つ:事後分布の近似やブートストラップを用いるThompsonサンプリングおよび関連アルゴリズムは、$1/\sqrt{n}$ スケール下で同じ弱極限を持つ。
  • 収束フレームワークは一般性に富み、ブートストラップベースの探索を含む他のサンプリングベースのバンディットアルゴリズムの解析に直接適用可能である。
  • 理論はマルチアームバンディットおよび有限または無限の行動集合、確率的・時間的に変化する行動集合を含む線形バンディットに適用可能である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。