[論文レビュー] Variational inference for the multi-armed contextual bandit
本稿では、文脈的マルチアームバンディットにおける変分トムソンサンプリング(VTS)を提案し、ガウス・ミックスチャーモデルを用いて複雑で未知の報酬分布を近似するための変分推論を用いる。真の報酬分布へのKLダイバージェンスを最小化するように変分パラメータを学習することで、モデルの誤り指定が生じても累積的リグレットが著しく低減され、制限的なパラメトリック仮定を伴う標準的トムソンサンプリングを上回る性能を発揮する。
In many biomedical, science, and engineering problems, one must sequentially decide which action to take next so as to maximize rewards. One general class of algorithms for optimizing interactions with the world, while simultaneously learning how the world operates, is the multi-armed bandit setting and, in particular, the contextual bandit case. In this setting, for each executed action, one observes rewards that are dependent on a given 'context', available at each interaction with the world. The Thompson sampling algorithm has recently been shown to enjoy provable optimality properties for this set of problems, and to perform well in real-world settings. It facilitates generative and interpretable modeling of the problem at hand. Nevertheless, the design and complexity of the model limit its application, since one must both sample from the distributions modeled and calculate their expected rewards. We here show how these limitations can be overcome using variational inference to approximate complex models, applying to the reinforcement learning case advances developed for the inference case in the machine learning community over the past two decades. We consider contextual multi-armed bandit applications where the true reward distribution is unknown and complex, which we approximate with a mixture model whose parameters are inferred via variational inference. We show how the proposed variational Thompson sampling approach is accurate in approximating the true distribution, and attains reduced regrets even with complex reward distributions. The proposed algorithm is valuable for practical scenarios where restrictive modeling assumptions are undesirable.
研究の動機と目的
- 文脈的バンディット設定において、制限的なパラメトリックモデルを超えたトムソンサンプリングの拡張を図ること。
- 真の分布が複雑または未知である場合の報酬モデリングにおけるモデルの誤り指定を解消すること。
- 混合モデルを用いて正確で解釈可能かつ柔軟な報酬分布のモデリングを可能にすること。
- 変分推論を用いて複雑な報酬構造を学習することで、不確実性下での逐次的意思決定における累積的リグレットを低減すること。
- 強いパラメトリック仮定を必要としない、スケーラブルで原理的整合性のある文脈的バンディットにおけるオンライン学習手法を提供すること。
提案手法
- 文脈的バンディットにおける複雑で未知の報酬分布を表現するために階層ベイジアン混合モデルを用いる。
- ガウス混合モデルのパラメータを学習するために、平均場近似を用いた変分推論を適用する。
- 真の後部確率分布と変分近似との間のKullback-Leiblerダイバージェンスを最小化することで、分布推定の正確性を保証する。
- 変分トムソンサンプリングを採用:変分後部分布からのサンプリングにより、探索と活用のバランスを図る。
- 証拠下限界(ELBO)最適化を用いて、変分パラメータの閉形式更新式を導出する。
- 新しい文脈-報酬ペアが到着する度に変分パラメータを段階的に更新することで、オンライン学習を統合する。
実験結果
リサーチクエスチョン
- RQ1変分推論は、指数型分布族に属さない複雑な報酬分布を文脈的バンディットで効果的に近似するために有効に用いられるか?
- RQ2モデルが誤って指定された場合でも、変分トムソンサンプリングは標準的トムソンサンプリングよりも累積的リグレットを低減するか?
- RQ3モデルの柔軟性(例えば、混合成分の数)が、複雑な報酬シナリオにおけるリグレットと推定精度に与える影響は何か?
- RQ4真の報酬分布が非対称的または重複している場合でも、提案手法は正確な期待報酬推定値を学習できるか?
- RQ5オンラインで逐次的意思決定を行う際の、モデルの複雑さとリグレット性能のトレードオフは何か?
主な発見
- 不均衡で重複する報酬混合を持つシナリオBにおいて、K=3成分のVTSはt=500でK=1のVTSと比較して累積的リグレットを40%低減した。
- K=3のVTSは真のモデル(K=2)と同等のリグレット性能を達成し、モデルの複雑さに対して高いロバストネスを示した。
- 複雑なシナリオにおいて、K=2およびK=3のVTSはK=1と比較して期待報酬推定のMSEが顕著に低く、学習精度の向上を示した。
- 単純なシナリオ(シナリオA)でさえも、K=1のVTSは良好な性能を示したが、より高いKのモデルの柔軟性が複雑さに起因する性能向上を可能にした。
- 変分近似はKLダイバージェンスを効果的に最小化し、正確な後部分布からのサンプリングと意思決定の改善を可能にした。
- 本手法は多様な報酬構造において低リグレットを維持し、制限的なパラメトリック仮定を必要とせず、優れた一般化性能を示した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。