Skip to main content
QUICK REVIEW

[Paper Review] Multiplayer Bandit Learning, from Competition to Cooperation

Simina Brânzei, Yuval Peres|arXiv (Cornell University)|Aug 3, 2019
Advanced Bandit Algorithms Research36 references6 citations
TL;DR

This paper studies multiplayer multi-armed bandit learning under varying cooperation levels, modeling competition (𝜆 = −1), neutrality (𝜆 = 0), and full cooperation (𝜆 = 1). It shows that competing and neutral players eventually coordinate on the same arm in every Nash equilibrium, while cooperating players may fail to settle due to strategic information hiding. Crucially, competing players explore less than a single player, whereas cooperating players explore more, with neutral players learning from each other to achieve higher rewards than solo play.

ABSTRACT

The stochastic multi-armed bandit model captures the tradeoff between exploration and exploitation. We study the effects of competition and cooperation on this tradeoff. Suppose there are $k$ arms and two players, Alice and Bob. In every round, each player pulls an arm, receives the resulting reward, and observes the choice of the other player but not their reward. Alice's utility is $Γ_A + λΓ_B$ (and similarly for Bob), where $Γ_A$ is Alice's total reward and $λ\in [-1, 1]$ is a cooperation parameter. At $λ= -1$ the players are competing in a zero-sum game, at $λ= 1$, they are fully cooperating, and at $λ= 0$, they are neutral: each player's utility is their own reward. The model is related to the economics literature on strategic experimentation, where usually players observe each other's rewards. With discount factor $β$, the Gittins index reduces the one-player problem to the comparison between a risky arm, with a prior $μ$, and a predictable arm, with success probability $p$. The value of $p$ where the player is indifferent between the arms is the Gittins index $g = g(μ,β) > m$, where $m$ is the mean of the risky arm. We show that competing players explore less than a single player: there is $p^* \in (m, g)$ so that for all $p > p^*$, the players stay at the predictable arm. However, the players are not myopic: they still explore for some $p > m$. On the other hand, cooperating players explore more than a single player. We also show that neutral players learn from each other, receiving strictly higher total rewards than they would playing alone, for all $ p\in (p^*, g)$, where $p^*$ is the threshold from the competing case. Finally, we show that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while this can fail for cooperating players.

Motivation & Objective

  • To understand how cooperation and competition affect exploration-exploitation tradeoffs in multiplayer stochastic bandit games.
  • To resolve a long-standing open question on whether players eventually settle on the same arm in Nash equilibria, particularly in competitive and neutral settings.
  • To quantify the value of information and its impact on exploration in zero-sum versus cooperative settings.
  • To compare equilibrium behavior and long-term rewards across different cooperation parameters 𝜆 ∈ [−1, 1].
  • To investigate whether neutral and competing players learn from each other and achieve higher rewards than solo play.

Proposed method

  • Models a two-player, two-arm bandit game with a known predictable arm (success probability 𝑝) and a risky arm with prior 𝜇.
  • Uses a cooperation parameter 𝜆 ∈ [−1, 1] to define player utilities as 𝑢𝑖 = Γ𝑖 + 𝜆Γ𝑗, interpolating between zero-sum (𝜆 = −1), neutral (𝜆 = 0), and fully cooperative (𝜆 = 1) games.
  • Analyzes Nash equilibria in both finite-horizon and discounted infinite-horizon settings, focusing on long-term behavior and equilibrium coordination.
  • Applies Gittins index theory to determine the threshold 𝑔(𝜇, 𝛽) where a single player is indifferent between arms.
  • Employs bounding techniques on expected rewards and net gains to analyze when players explore or stay put, especially in the limit as 𝛽 → 1.
  • Uses strategy construction and payoff comparison (e.g., Bob copying Alice’s past actions) to derive lower bounds on net gains and infer equilibrium behavior.

Experimental results

Research questions

  • RQ1Do competing and neutral players always eventually settle on the same arm in every Nash equilibrium?
  • RQ2Is exploration reduced in zero-sum games compared to single-player bandit settings?
  • RQ3Do neutral players learn from each other, achieving strictly higher rewards than solo play in equilibrium?
  • RQ4Can cooperating players fail to settle on a single arm, even in equilibrium?
  • RQ5How do the thresholds 𝑝∗ and 𝑒𝑝 relate to the Gittins index 𝑔, and are they monotonic in 𝛽 and 𝜆?

Key findings

  • In every Nash equilibrium, competing and neutral players eventually settle on the same arm, even if it is not optimal, while this coordination fails for cooperating players.
  • Competing players explore less than a single player: there exists 𝑝∗ ∈ (𝑚, 𝑔) such that for all 𝑝 > 𝑝∗, players remain at the predictable arm in all equilibria.
  • Despite reduced exploration, competing players still explore for all 𝑝 > 𝑚, showing they are not myopic.
  • Cooperating players (𝜆 = 1) explore more than a single player, with exploration increasing relative to the single-agent optimum.
  • Neutral players learn from each other: in every perfect Bayesian equilibrium, each player receives strictly higher expected total reward than when playing alone, for all 𝑝 ∈ (𝑝∗, 𝑔).
  • For 𝑝 < 5/9, competing players explore the risky arm in some equilibria; for 𝑝 > 2 − √2 ≈ 0.586, they do not explore the risky arm in any equilibrium.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.