Skip to main content
QUICK REVIEW

[論文レビュー] Restless Bandits with Many Arms: Beating the Central Limit Theorem

Xiangyu Zhang, Peter I. Frazier|arXiv (Cornell University)|Jul 25, 2021
Advanced Bandit Algorithms Research参考文献 37被引用数 9
ひとこと要約

本稿は、多数の腕を有する restless bandits における fluid-priority 政策を導入し、非退化性条件の下で最適性ギャップが O(1) であることを証明する。これは中心極限定理が示唆する古典的 O(√N) 界を上回る。この手法は、大規模なマルコフ決定過程における広範なインデックス政策クラスのより緊密な収束速度を確立するために、フロイド的および拡散スケールの極限を利用する。

ABSTRACT

We consider finite-horizon restless bandits with multiple pulls per period, which play an important role in recommender systems, active learning, revenue management, and many other areas. While an optimal policy can be computed, in principle, using dynamic programming, the computation required scales exponentially in the number of arms $N$. Thus, there is substantial value in understanding the performance of index policies and other policies that can be computed efficiently for large $N$. We study the growth of the optimality gap, i.e., the loss in expected performance compared to an optimal policy, for such policies in a classical asymptotic regime proposed by Whittle in which $N$ grows while holding constant the fraction of arms that can be pulled per period. Intuition from the Central Limit Theorem and the tightest previous theoretical bounds suggest that this optimality gap should grow like $O(\sqrt{N})$. Surprisingly, we show that it is possible to outperform this bound. We characterize a non-degeneracy condition and a wide class of novel practically-computable policies, called fluid-priority policies, in which the optimality gap is $O(1)$. These include most widely-used index policies. When this non-degeneracy condition does not hold, we show that fluid-priority policies nevertheless have an optimality gap that is $O(\sqrt{N})$, significantly generalizing the class of policies for which convergence rates are known. We demonstrate that fluid-priority policies offer state-of-the-art performance on a collection of restless bandit problems in numerical experiments.

研究の動機と目的

  • 大規模な restless bandit 問題において、シミュレーション結果が定数の最適性ギャップを示すのに対し、理論的境界が O(√N) の増加を示すというギャップを解消すること。
  • 腕の数 N が増加する非漸近的状況において、予算が比例的に増加する条件下で、インデックス政策の性能を分析する一般枠組みを構築すること。
  • fluid-priority 政策が O(1) の最適性ギャップを達成する非退化性条件を同定すること。これは、従来の O(√N) 界に比べて顕著に改善される。
  • ほとんどの広く用いられるインデックス政策が計算可能であるという点を踏まえ、実用的に計算可能な政策の広いクラスをカバーする理論的分析を拡張すること。
  • 多様な restless bandit 問題における数値実験を通じて、fluid-priority 政策が最先端の性能を達成することを示すこと。

提案手法

  • Brown ら (2020) や Hu と Frazier (2017) のインデックス政策の主要な特徴を一般化する fluid-priority 政策という一般クラスを導入する。
  • フロイド一貫性を定義する:N → ∞ のとき、政策の正規化された状態および行動回数が、確率的にフロイド極限 z_t および x_t に収束する。
  • フロイド一貫性が o(N) の最適性ギャップを示すことを確立する。これは、より緊密な境界への基礎的ステップである。
  • 拡散正則性の概念を導入する:政策が誘導する写像はリプシッツ連続で、ゼロで有界で、N → ∞ のとき収束する必要がある。
  • 拡散正則性が、拡散スケール統計量 (Z̃_t^N, X̃_t^N) の2次モーメントの一様有界性およびその分布収束を保証することを証明する。
  • 拡散スケールの偏差の一様有界性を用いて、より弱い条件下でも、最適値と政策値のギャップが O(√N) であることを示す。

実験結果

リサーチクエスチョン

  • RQ1特定の条件下では、restless bandits における古典的 O(√N) の最適性ギャップ境界を改善できるか?
  • RQ2多くの腕を有する漸近的状況において、どの政策クラスが O(1) の最適性ギャップを達成できるか? また、そのために必要な条件は何か?
  • RQ3フロイド的および拡散スケールの極限を用いて、大規模な restless bandit 問題におけるインデックス政策の収束速度をどのように分析できるか?
  • RQ4既存のインデックス政策(例:Brown ら 2020、Hu と Frazier 2017)は、より緊密な性能境界を得るための新しい正則性条件をどの程度満たしているか?
  • RQ5fluid-priority 政策が、N にかかわらず常に定数の最適性ギャップを達成する非退化性条件が存在するか?

主な発見

  • 非退化性条件の下で、fluid-priority 政策は O(1) の最適性ギャップを達成する。これは、中心極限定理が示唆する O(√N) 界を厳密に上回る。
  • 政策のフロイド一貫性は、広い範囲の政策に適用可能な o(N) の最適性ギャップを示す。これは一般的な結果である。
  • フロイド一貫性よりも強い条件である拡散正則性は、O(√N) の最適性ギャップを保証する。これは、従来の結果をより広い政策クラスに一般化する。
  • Brown ら (2020) や Hu と Frazier (2017) が提唱した政策が拡散正則性を満たすことが示され、新しい枠組みによりそれらの O(√N) の最適性ギャップが確認された。
  • 数値実験により、fluid-priority 政策が多様な restless bandit 問題において最先端の性能を達成することが示された。
  • 理論的枠組みにより、従来の限界を超えるより緊密な境界を証明する道筋が提供され、特に CLT に基づく直観が失敗する状況において顕著である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。