[論文レビュー] PAC Identification of Many Good Arms in Stochastic Multi-Armed Bandits
本稿は、確率的マルチアームバンディットにおける$k$番目の最良の$m$本の腕から選ぶPAC同定問題を導入し、最良のサブセット選択と単一腕同定を一般化する。LUCB-$k$-$m$アルゴリズムを提案し、定数要因の違いを除いて既知の下界に一致するサンプル複雑性を達成する。また、対数的ギャップが最適性に近い無限腕バンディットへとフレームワークを拡張する。
We consider the problem of identifying any $k$ out of the best $m$ arms in an $n$-armed stochastic multi-armed bandit. Framed in the PAC setting, this particular problem generalises both the problem of `best subset selection' and that of selecting `one out of the best m' arms [arcsk 2017]. In applications such as crowd-sourcing and drug-designing, identifying a single good solution is often not sufficient. Moreover, finding the best subset might be hard due to the presence of many indistinguishably close solutions. Our generalisation of identifying exactly $k$ arms out of the best $m$, where $1 \leq k \leq m$, serves as a more effective alternative. We present a lower bound on the worst-case sample complexity for general $k$, and a fully sequential PAC algorithm, \GLUCB, which is more sample-efficient on easy instances. Also, extending our analysis to infinite-armed bandits, we present a PAC algorithm that is independent of $n$, which identifies an arm from the best $ρ$ fraction of arms using at most an additive poly-log number of samples than compared to the lower bound, thereby improving over [arcsk 2017] and [Aziz+AKA:2018]. The problem of identifying $k > 1$ distinct arms from the best $ρ$ fraction is not always well-defined; for a special class of this problem, we present lower and upper bounds. Finally, through a reduction, we establish a relation between upper bounds for the `one out of the best $ρ$' problem for infinite instances and the `one out of the best $m$' problem for finite instances. We conjecture that it is more efficient to solve `small' finite instances using the latter formulation, rather than going through the former.
研究の動機と目的
- 単一の最適な腕を同定するのでは不十分な実用的状況、例えば分散型コロナサージャーやドラッグ設計において、複数の優れた解決策を必要とする。
- 既存の問題を一般化する:最良のサブセット選択($k=m$)と最良の$m$本から1本を選ぶ($k=1$)問題を、$n$本の腕の中から上位$m$本のうち$k$本の腕を同定する統一されたフレームワークに統合する。
- 有限および無限腕バンディット設定における一般化された$(k,m,n)$問題の理論的サンプル複雑性境界を確立する。
- 最小限のサンプリングで$k$個の良い腕を効率的に同定できる、完全に逐次的かつ適応的なPACアルゴリズム、LUCB-$k$-$m$を設計する。
- 無限腕バンディットにおける$k$個の異なる$[\epsilon,\rho]$最適腕の同定可能性と複雑性を分析し、有限と無限のインスタンスを還元によって関連付ける。
提案手法
- 新しい問題定式化を提案:確率的マルチアームバンディットにおいて、$n$本の腕の中から上位$m$本のうちの$k$本の腕をPAC設定で同定する。
- 上界と下界の信頼区間を用いて、逐次的に非最適な腕を除外する完全に逐次的かつ適応的なLUCB-$k$-$m$アルゴリズムを導入する。
- $(k,m,n)$問題における最悪ケースのサンプル複雑性の下界を導出する。これは、$k=1$および$k=m$の場合の既存の下界を一般化する。
- 無限腕バンディットへの分析を拡張する際、$\rho$-最適腕(平均値の$\rho$-分位数から$\epsilon$以内の腕)を定義し、下界に加法的$O(\log^2(1/\delta))$の誤差で収まるPACアルゴリズムを提示する。
- $k$-out-of-best-$\rho$-fraction問題のwell-posedなインスタンスのクラスを特定し、理論的保証を伴う対応するアルゴリズムを提供する。
- 無限腕$(k,\rho)$-問題と有限腕$(k,m,n)$-問題の間の還元を確立し、有限バージョンを解く方がより効率的である可能性を示唆する。
実験結果
リサーチクエスチョン
- RQ1確率的バンディットにおいて、$n$本の腕の中から上位$m$本のうちの$k$本の腕を同定する際の、根本的なサンプル複雑性の下界は何か?
- RQ2一般化された$(k,m,n)$問題に対して、近似的に最適なサンプル複雑性を達成できる完全に逐次的かつ適応的なアルゴリズムを設計できるか?
- RQ3無限腕バンディットにおいて、$k$本の腕を最良の$\rho$分率から同定するサンプル複雑性はどのようにスケーリングされるか?また、情報理論的下界に近づけることができるか?
- RQ4$k$個の異なる$[\epsilon,\rho]$最適腕を同定する問題がwell-posedである条件は何か?そして、効率的に解けるか?
- RQ5サンプル効率の観点から、無限腕$(k,\rho)$-問題を解くのと比較して、有限腕$(k,m,n)$-問題を解く際に構造的な利点があるか?
主な発見
- LUCB-$k$-$m$アルゴリズムは、$k=1$および$k=m$の場合に既知の下界に定数要因の違いを除いて一致する期待サンプル複雑性を達成する。
- 実験的評価では、$n$が増加するにつれて、$\mathcal{F}_2$などの既存のアルゴリズムに比べてLUCB-$k$-$m$が顕著にサンプル効率に優れている。
- 無限腕バンディットにおいて、提案されたアルゴリズムは$[\epsilon,\rho]$最適腕を、下界に加法的$O(\log^2(1/\delta))$の誤差で同定するサンプル複雑性を達成し、先行研究を改善する。
- $k$個の異なる$[\epsilon,\rho]$最適腕を同定する問題は常にwell-posedではないが、特定のインスタンスクラスではwell-posedな定式化が可能であり、理論的境界が保証される。
- 無限腕$(k,\rho)$-問題と有限腕$(k,m,n)$-問題の間の還元が確立され、有限インスタンスの方が効率的に解ける可能性を示唆する。
- 著者らは、有限腕$(k,m,n)$-問題を解く方が、対応する無限腕$(k,\rho)$-問題を解くよりもサンプル効率に優れていると仮説を立てており、今後の研究の方向性として残されている。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。