Skip to main content
QUICK REVIEW

[論文レビュー] Bayesian Exploration: Incentivizing Exploration in Bayesian Games

Yishay Mansour, Aleksandrs Slivkins|arXiv (Cornell University)|Feb 24, 2016
Game Theory and Applications被引用数 8
ひとこと要約

この論文は、報酬の移転なしに、ベイジアンゲームにおけるエージェントが不確実な行動を探索するようにインcentivizeするためのフレームワーク、Bayesian Explorationを導入する。『探索可能な行動』を特定し、インcentive-compatibleな推薦ポリシーを用いることで、決定的設定では定数レグレット、確率的設定では対数レグレットを達成し、単一エージェント探索に関する先行研究を著しく改善する。

ABSTRACT

We consider a ubiquitous scenario in the Internet economy when individual decision-makers (henceforth, agents) both produce and consume information as they make strategic choices in an uncertain environment. This creates a three-way tradeoff between exploration (trying out insufficiently explored alternatives to help others in the future), exploitation (making optimal decisions given the information discovered by other agents), and incentives of the agents (who are myopically interested in exploitation, while preferring the others to explore). We posit a principal who controls the flow of information from agents that came before, and strives to coordinate the agents towards a socially optimal balance between exploration and exploitation, not using any monetary transfers. The goal is to design a recommendation policy for the principal which respects agents' incentives and minimizes a suitable notion of regret. We extend prior work in this direction to allow the agents to interact with one another in a shared environment: at each time step, multiple agents arrive to play a Bayesian game, receive recommendations, choose their actions, receive their payoffs, and then leave the game forever. The agents now face two sources of uncertainty: the actions of the other agents and the parameters of the uncertain game environment. Our main contribution is to show that the principal can achieve constant regret when the utilities are deterministic (where the constant depends on the prior distribution, but not on the time horizon), and logarithmic regret when the utilities are stochastic. As a key technical tool, we introduce the concept of explorable actions, the actions which some incentive-compatible policy can recommend with non-zero probability. We show how the principal can identify (and explore) all explorable actions, and use the revealed information to perform optimally.

研究の動機と目的

  • 自己中心的で短期的思考のエージェントが参加するベイジアンゲームにおいて、集団的探索が社会に利益をもたらすが、そのインセンティブを提供する課題に取り組む。
  • 報酬の移転なしに、エージェントのインセンティブを尊重する(ベイジアンインセンティブコンプライアンスを介して)レコメンデーションポリシーを設計し、レグレットを最小限に抑える。
  • 単一エージェント探索モデルを、エージェントが共有の不確実な環境で相互作用する多エージェント設定に拡張する。
  • 行動が『探索可能』である条件を特定し、インcentive制約のもとでそれを効率的に特定・探索する方法を示す。
  • 任意のプラインシパルの目的関数を想定しても、決定的設定では定数、確率的設定では対数の最適なレグレットバウンドを達成する。

提案手法

  • 『探索可能な行動』—あるインセンティブコンプライアンスポリシーのもとで正の確率で推薦可能な行動—という概念を導入する。
  • すべての探索可能な行動を特定・探索する、最大限に探索を促進するサブルーチンを開発し、BIC準拠のレコメンデーションメカニズムを用いる。
  • 複数のエージェントが各ラウンドに到着し、ベイジアンゲームをプレーし、レコメンデーションを受け、行動をとり、退場するという繰り返しゲームの枠組みを採用する。
  • 全体のポリシーがベイジアンインセンティブコンプライアンスを保ちつつ、すべての探索可能な行動を探索できるように、BICサブルーチンの合成を用いる。
  • 確率的報酬設定を扱うために、期待報酬とシグナル設計の近似技術を適用する。
  • 探索と活用を分離しつつ、ラウンド間でインセンティブコンプライアンスを維持する、新しい分析フレームワークに依存する。

実験結果

リサーチクエスチョン

  • RQ1プラインシパルは、報酬の移転なしに、多エージェントのベイジアンゲーム設定においてエージェントが探索を促せるレコメンデーションポリシーを設計できるか?
  • RQ2自己中心的エージェントと不完全情報のもとで、ゲーム理論的設定において行動が『探索可能』とされる条件は何か?
  • RQ3インセンティブコンプライアンスを保証しつつ、決定的報酬設定で定数レグレットを達成するにはどうすればよいか?
  • RQ4プラインシパルが完全なBICではなくδ-BICポリシーと競合する確率的報酬設定では、レグレットの根本的限界は何か?
  • RQ5時間的に変化する文脈や要約的ゲーム表現を含む設定へ、このフレームワークを拡張可能か?

主な発見

  • 決定的報酬設定では、プラインシパルは定数レグレットを達成でき、その定数は時間の長さに依存せず、事前分布にのみ依存する。
  • 確率的報酬設定では、任意のδ > 0に対してδ-BICポリシーと競合する際、プラインシパルは対数レグレットを達成する。
  • すべての探索可能な行動—BICのもとで正の確率で推薦可能な行動—は、計算的に効率的なサブルーチンを用いて特定・探索可能である。
  • このフレームワークは、プラインシパルの目的関数がエージェントの累積報酬と一致する必要がないため、目的設計の柔軟性を提供する。
  • 単一エージェント設定で必須とされていた『すべての行動が探索可能』という仮定を排除したことで、先行研究を著しく改善している。
  • 分析により、完全なBICポリシー(δ = 0)と競合する場合、データ依存のサンプリング要件がBICの合成を破壊するため、未解決の課題であることが明らかになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。