Skip to main content
QUICK REVIEW

[논문 리뷰] Multiplayer Bandit Learning, from Competition to Cooperation

Simina Brânzei, Yuval Peres|arXiv (Cornell University)|2019. 08. 03.
Advanced Bandit Algorithms Research참고 문헌 36인용 수 6
한 줄 요약

이 논문은 협력 수준이 다양할 때 다중 플레이어 다손대기 문제 학습을 연구하며, 경쟁(𝜆 = −1), 중립(𝜆 = 0), 완전 협력(𝜆 = 1)을 모델링한다. 경쟁 및 중립 플레이어는 모든 내시 균형에서 결국 같은 손을 선택하는 반면, 협력 플레이어는 전략적 정보 은폐로 인해 정착하지 못할 수 있다. 핵심적으로, 경쟁 플레이어는 단일 플레이어보다 탐색을 적게 하며, 협력 플레이어는 더 많이 탐색한다. 중립 플레이어는 상호 학습을 통해 단독 플레이보다 더 높은 보상을 달 đạt한다.

ABSTRACT

The stochastic multi-armed bandit model captures the tradeoff between exploration and exploitation. We study the effects of competition and cooperation on this tradeoff. Suppose there are $k$ arms and two players, Alice and Bob. In every round, each player pulls an arm, receives the resulting reward, and observes the choice of the other player but not their reward. Alice's utility is $Γ_A + λΓ_B$ (and similarly for Bob), where $Γ_A$ is Alice's total reward and $λ\in [-1, 1]$ is a cooperation parameter. At $λ= -1$ the players are competing in a zero-sum game, at $λ= 1$, they are fully cooperating, and at $λ= 0$, they are neutral: each player's utility is their own reward. The model is related to the economics literature on strategic experimentation, where usually players observe each other's rewards. With discount factor $β$, the Gittins index reduces the one-player problem to the comparison between a risky arm, with a prior $μ$, and a predictable arm, with success probability $p$. The value of $p$ where the player is indifferent between the arms is the Gittins index $g = g(μ,β) > m$, where $m$ is the mean of the risky arm. We show that competing players explore less than a single player: there is $p^* \in (m, g)$ so that for all $p > p^*$, the players stay at the predictable arm. However, the players are not myopic: they still explore for some $p > m$. On the other hand, cooperating players explore more than a single player. We also show that neutral players learn from each other, receiving strictly higher total rewards than they would playing alone, for all $ p\in (p^*, g)$, where $p^*$ is the threshold from the competing case. Finally, we show that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while this can fail for cooperating players.

연구 동기 및 목표

  • 다중 플레이어 스토하스틱 밴딧 게임에서 협력과 경쟁이 탐색-이용 트레이드오프에 미치는 영향을 이해하기 위해.
  • 경쟁 및 중립 설정에서 플레이어가 결국 동일한 손을 선택하는지 여부에 대한 오랫동안 미해결된 열린 문제를 해결하기 위해.
  • 정보의 가치와 그 영향을 제로섬 대비 협력 설정에서의 탐색에 정량화하기 위해.
  • 협력 매개변수 𝜆 ∈ [−1, 1]에 따라 균형 행동과 장기 보상의 비교를 위해.
  • 중립 및 경쟁 플레이어가 서로 학습하여 단독 플레이보다 더 높은 보상을 달 đạt하는지 조사하기 위해.

제안 방법

  • 알고리즘의 성능을 평가하기 위해, 알려진 예측 가능한 손(성공 확률 𝑝)과 사전 분포 𝜇를 가진 위험한 손을 가진 두 플레이어, 두 손 밴딧 게임을 모델링한다.
  • 플레이어의 보상 함수를 𝑢𝑖 = Γ𝑖 + 𝜆Γ𝑗로 정의하여 협력 매개변수 𝜆 ∈ [−1, 1]을 사용하며, 이는 제로섬(𝜆 = −1), 중립(𝜆 = 0), 완전 협력(𝜆 = 1) 게임을 보간한다.
  • 유한 기간 및 할인 무한 기간 설정에서 내시 균형을 분석하며, 장기적 행동과 균형 정착에 집중한다.
  • 단일 플레이어가 두 손 사이에서 무관심이 되는 임계값 𝑔(𝜇, 𝛽)를 결정하기 위해 기티너스 인덱스 이론을 적용한다.
  • 기대 보상과 순수익에 대한 경계 기법을 사용하여, 특히 𝛽 → 1일 때 탐색 또는 정착 여부를 분석한다.
  • 전략 구성 및 보상 비교(예: 박브가 앨리스의 과거 행동을 모방)를 통해 순수익의 하한을 유도하고 균형 행동을 추론한다.

실험 결과

연구 질문

  • RQ1경쟁 및 중립 플레이어는 모든 내시 균형에서 결국 동일한 손을 선택하는가?
  • RQ2제로섬 게임에서는 단일 플레이어 밴딧 설정보다 탐색이 줄어드는가?
  • RQ3중립 플레이어는 서로 학습하여 단독 플레이보다 엄밀히 더 높은 보상을 달달하는가?
  • RQ4협력 플레이어는 조차도 균형 상태에서도 단일 손에 정착하지 못할 수 있는가?
  • RQ5임계값 𝑝∗ 및 𝑒𝑝는 기티너스 인덱스 𝑔와 어떻게 관련되어 있으며, 𝛽 및 𝜆에 대해 단조 증가하는가?

주요 결과

  • 모든 내시 균형에서 경쟁 및 중립 플레이어는 결국 동일한 손을 선택한다. 이는 최적의 손이 아니더라도 그렇다. 반면 협력 플레이어는 정착에 실패할 수 있다.
  • 경쟁 플레이어는 단일 플레이어보다 탐색을 적게 한다. 모든 균형에서 예측 가능한 손에 머무르는 𝑝∗ ∈ (𝑚, 𝑔)가 존재하며, 𝑝 > 𝑝∗일 경우에 해당한다.
  • 다시 말해, 탐색 감소에도 불구하고 경쟁 플레이어는 모든 𝑝 > 𝑚에서 탐색을 수행하므로, 이는 이성적이지 않다는 것을 보여준다.
  • 협력 플레이어(𝜆 = 1)는 단일 플레이어보다 더 많은 탐색을 하며, 단일 에이전트 최적화 대비 탐색이 증가한다.
  • 중립 플레이어는 상호 학습을 한다. 모든 완전 정보 베이지안 균형에서, 모든 𝑝 ∈ (𝑝∗, 𝑔)에 대해 각 플레이어는 단독 플레이 시보다 엄밀히 더 높은 기대 총보상을 얻는다.
  • 𝑝 < 5/9일 경우, 경쟁 플레이어는 일부 균형에서 위험한 손을 탐색한다. 반면 𝑝 > 2 − √2 ≈ 0.586일 경우, 어떤 균형에서도 위험한 손을 탐색하지 않는다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.