[論文レビュー] On Worst-case Regret of Linear Thompson Sampling.
本稿では、線形トムソンサンプリング(LinTS)が事後分散の拡張なしでは最悪ケースにおいて線形のリグレットを被ることを確立し、サブリニアリグレットを達成するためには $\tilde{\mathcal{O}}(\sqrt{d})$ の拡張が必須であることを証明している。さらに、やや弱い条件下では、はるかに小さい $\tilde{\mathcal{O}}(1)$ の拡張で十分であることが示され、LinTSの最悪ケースリグレットバウンドに関する未解決問題が解決された。
In this paper, we consider the worst-case regret of Linear Thompson Sampling (LinTS) for the linear bandit problem. \citet{russo2014learning} show that the Bayesian regret of LinTS is bounded above by $\widetilde{\mathcal{O}}(d\sqrt{T})$ where $T$ is the time horizon and $d$ is the number of parameters. While this bound matches the minimax lower-bounds for this problem up to logarithmic factors, the existence of a similar worst-case regret bound is still unknown. The only known worst-case regret bound for LinTS, due to \cite{agrawal2013thompson,abeille2017linear}, is $\widetilde{\mathcal{O}}(d\sqrt{dT})$ which requires the posterior variance to be inflated by a factor of $\widetilde{\mathcal{O}}(\sqrt{d})$. While this bound is far from the minimax optimal rate by a factor of $\sqrt{d}$, in this paper we show that it is the best possible one can get, settling an open problem stated in \cite{russo2018tutorial}. Specifically, we construct examples to show that, without the inflation, LinTS can incur linear regret up to time $\exp(\Omega(d))$. We then demonstrate that, under mild conditions, a slightly modified version of LinTS requires only an $\widetilde{\mathcal{O}}(1)$ inflation where the constant depends on the diversity of the optimal arm.
研究の動機と目的
- 事後分散の過剰な拡張なしにLinTSがミニマックス最適な最悪ケースリグレットを達成できるかどうかという未解決問題を解明すること。
- 敵対的設定において線形リグレットを回避するためのLinTSに必要な最小の事後分散拡張量を特定すること。
- 最適腕の構造的仮定がやや弱い条件下で、定数の拡張係数で十分である条件を同定すること。
- 拡張なしの場合に線形リグレットが生じることを示す明示的な例を構築すること。
提案手法
- 事後分散を拡張しない場合にLinTSが時間 $\exp(\Omega(d))$ まで線形リグレットを被る敵対的線形バンディットインスタンスを構築する。
- 最適腕の多様性に依存するリグレットの依存関係を分析し、必要な拡張係数を制限する。
- 最適腕にやや弱い多様性条件が成り立つ場合に、$\tilde{\mathcal{O}}(1)$ の拡張係数を使用する修正されたLinTSアルゴリズムを導入する。
- 集中不等式と事後分散解析を用いて最悪ケースのリグレットバウンドを導出する。
- Hardなインスタンスを構築することで、必要な拡張係数の下界を確立する。
実験結果
リサーチクエスチョン
- RQ1事後分散の拡張なしに、LinTSがサブリニアな最悪ケースリグレットを達成することは可能か?
- RQ2最悪ケースにおいて線形リグレットを回避するためのLinTSに必要な最小の拡張係数は何か?
- RQ3最適腕にやや弱い構造的仮定が成り立つ場合、必要な拡張係数を $\tilde{\mathcal{O}}(1)$ にまで低減できるか?
- RQ4最適腕の多様性は、必要な拡張係数にどのように影響するか?
主な発見
- 事後分散の拡張なしでは、LinTSは時間 $\exp(\Omega(d))$ まで線形リグレットを被る可能性があり、最悪ケースにおいて $\tilde{\mathcal{O}}(\sqrt{d})$ の拡張が必須であることを示している。
- 最悪ケースのリグレットが $\tilde{\mathcal{O}}(d\sqrt{T})$ に抑えられるのは、事後分散が $\tilde{\mathcal{O}}(\sqrt{d})$ だけ拡張された場合に限る。
- 最適腕の多様性にやや弱い条件が成り立つ場合、$\tilde{\mathcal{O}}(1)$ の拡張係数で十分であり、$\tilde{\mathcal{O}}(d\sqrt{T})$ のリグレットを達成できる。
- 本稿では、\cite{russo2018tutorial} で提起された未解決問題を解決し、LinTSの最悪ケースバウンドとして $\tilde{\mathcal{O}}(\sqrt{d})$ が最良であることを示した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。