Skip to main content
QUICK REVIEW

[論文レビュー] Multiplayer Bandit Learning, from Competition to Cooperation

Simina Brânzei, Yuval Peres|arXiv (Cornell University)|Aug 3, 2019
Advanced Bandit Algorithms Research参考文献 36被引用数 6
ひとこと要約

この論文は、協力度が異なる状況下におけるマルチプレイヤー多腕バンディット学習を研究しており、競合(𝜆 = −1)、中立(𝜆 = 0)、完全協力(𝜆 = 1)の状況をモデル化している。競合および中立プレイヤーは、すべてのナッシュ均衡において最終的に同じ腕に到達することが示され、協力的プレイヤーは戦略的情報隠蔽のため、安定化に失敗する可能性がある。重要ポイントとして、競合プレイヤーは単一プレイヤー未満の探索を行うが、協力プレイヤーは単一プレイヤーを上回る探索を行う。中立プレイヤーは互いに学習することで、単独プレイよりも高い報酬を達成する。

ABSTRACT

The stochastic multi-armed bandit model captures the tradeoff between exploration and exploitation. We study the effects of competition and cooperation on this tradeoff. Suppose there are $k$ arms and two players, Alice and Bob. In every round, each player pulls an arm, receives the resulting reward, and observes the choice of the other player but not their reward. Alice's utility is $Γ_A + λΓ_B$ (and similarly for Bob), where $Γ_A$ is Alice's total reward and $λ\in [-1, 1]$ is a cooperation parameter. At $λ= -1$ the players are competing in a zero-sum game, at $λ= 1$, they are fully cooperating, and at $λ= 0$, they are neutral: each player's utility is their own reward. The model is related to the economics literature on strategic experimentation, where usually players observe each other's rewards. With discount factor $β$, the Gittins index reduces the one-player problem to the comparison between a risky arm, with a prior $μ$, and a predictable arm, with success probability $p$. The value of $p$ where the player is indifferent between the arms is the Gittins index $g = g(μ,β) > m$, where $m$ is the mean of the risky arm. We show that competing players explore less than a single player: there is $p^* \in (m, g)$ so that for all $p > p^*$, the players stay at the predictable arm. However, the players are not myopic: they still explore for some $p > m$. On the other hand, cooperating players explore more than a single player. We also show that neutral players learn from each other, receiving strictly higher total rewards than they would playing alone, for all $ p\in (p^*, g)$, where $p^*$ is the threshold from the competing case. Finally, we show that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while this can fail for cooperating players.

研究の動機と目的

  • 協力と競合がマルチプレイヤー確率的バンディットゲームにおける探索と活用のトレードオフに与える影響を理解すること。
  • 長年の未解決課題である、特に競合および中立設定下で、プレイヤーがナッシュ均衡において最終的に同じ腕に到達するかどうかを解明すること。
  • 情報の価値とその探索への影響を、ゼロサム対比と協力的状況で定量化すること。
  • 異なる協力パラメータ 𝜆 ∈ [−1, 1] における均衡行動と長期的報酬を比較すること。
  • 中立および競合プレイヤーが互いに学習し、単独プレイよりも高い報酬を均衡で達成するかどうかを調査すること。

提案手法

  • 既知の予測可能な腕(成功確率 𝑝)と事前分布 𝜇 があるリスクの高い腕を持つ2プレイヤー2腕のバンディットゲームをモデル化する。
  • 協力パラメータ 𝜆 ∈ [−1, 1] を用いて、プレイヤーの報酬関数を 𝑢𝑖 = Γ𝑖 + 𝜆Γ𝑗 として定義し、ゼロサム(𝜆 = −1)、中立(𝜆 = 0)、完全協力(𝜆 = 1)の状況を補間する。
  • 有限時限および割引無限時限の両設定におけるナッシュ均衡を分析し、長期的行動と均衡の協調性に注目する。
  • ギッティンズ指数理論を用いて、単一プレイヤーが両腕に対して無差別になる閾値 𝑔(𝜇, 𝛽) を特定する。
  • 期待報酬と純利益のバウンディング技術を用い、特に 𝛽 → 1 の極限において、プレイヤーが探索するか停止するかを分析する。
  • 戦略構築と報酬比較(例:ボブがアリスの過去の行動を模倣)を用いて、純利益の下界を導出し、均衡行動を推論する。

実験結果

リサーチクエスチョン

  • RQ1競合および中立プレイヤーは、すべてのナッシュ均衡において、最終的に同じ腕に到達するのか?
  • RQ2ゼロサムゲームでは、単一プレイヤーのバンディット設定に比べ、探索が減少するのか?
  • RQ3中立プレイヤーは互いに学習し、均衡において単独プレイよりも厳密に高い報酬を達成するのか?
  • RQ4協力的プレイヤー(𝜆 = 1)でさえ、均衡においても単一の腕に到達しないことがあるのか?
  • RQ5閾値 𝑝∗ と 𝑒𝑝 はギッティンズ指数 𝑔 とどのように関係し、𝛽 および 𝜆 に関して単調性を示すのか?

主な発見

  • すべてのナッシュ均衡において、競合および中立プレイヤーは最終的に同じ腕に到達するが、最適ではない場合でも、協力的プレイヤーではこの協調が失敗する。
  • 競合プレイヤーは単一プレイヤー未満の探索を行う:すべての均衡で予測可能な腕に留まるような 𝑝∗ ∈ (𝑚, 𝑔) が存在し、すべての 𝑝 > 𝑝∗ に対して成立する。
  • 減少した探索にもかかわらず、競合プレイヤーはすべての 𝑝 > 𝑚 に対して探索を継続するため、短視眼的ではないことが示される。
  • 協力的プレイヤー(𝜆 = 1)は、単一プレイヤーを上回る探索を行う。特に、単一エージェントの最適解に比べて、探索が増加する。
  • 中立プレイヤーは互いに学習する:すべての 𝑝 ∈ (𝑝∗, 𝑔) に対して、すべての完全ベイジアン均衡において、各プレイヤーは単独プレイ時よりも厳密に高い期待総報酬を得る。
  • 𝑝 < 5/9 の場合、競合プレイヤーは一部の均衡でリスクの高い腕を探索するが、𝑝 > 2 − √2 ≈ 0.586 の場合、あらゆる均衡でリスクの高い腕を探索しない。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。