Skip to main content
QUICK REVIEW

[論文レビュー] Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games

Xiangyu Liu, Hangtian Jia|arXiv (Cornell University)|Jun 9, 2021
Experimental Behavioral Economics Studies参考文献 3被引用数 9
ひとこと要約

本稿は、ゼロ和マルコフゲームにおけるオープンエンドドラーニングにおいて、行動多様性(BD)と応答多様性(RD)を測定・促進する統一枠組みを提案する。BDは占有測度の差異として定義され、RDはゲームスケープの凸包への幾何的射影を通じて定義される。この手法により、ポリシー集団の多様性が向上し、行列ゲームでは低い特徴的脆弱性と高い集団効果性が達成され、Google Research Footballでも優れた性能を示す。

ABSTRACT

Measuring and promoting policy diversity is critical for solving games with strong non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). With that in mind, maintaining a pool of diverse policies via open-ended learning is an attractive solution, which can generate auto-curricula to avoid being exploited. However, in conventional open-ended learning algorithms, there are no widely accepted definitions for diversity, making it hard to construct and evaluate the diverse policies. In this work, we summarize previous concepts of diversity and work towards offering a unified measure of diversity in multi-agent open-ended learning to include all elements in Markov games, based on both Behavioral Diversity (BD) and Response Diversity (RD). At the trajectory distribution level, we re-define BD in the state-action space as the discrepancies of occupancy measures. For the reward dynamics, we propose RD to characterize diversity through the responses of policies when encountering different opponents. We also show that many current diversity measures fall in one of the categories of BD or RD but not both. With this unified diversity measure, we design the corresponding diversity-promoting objective and population effectivity when seeking the best responses in open-ended learning. We validate our methods in both relatively simple games like matrix game, non-transitive mixture model, and the complex extit{Google Research Football} environment. The population found by our methods reveals the lowest exploitability, highest population effectivity in matrix game and non-transitive mixture model, as well as the largest goal difference when interacting with opponents of various levels in extit{Google Research Football}.

研究の動機と目的

  • オープンエンドドラーニングにおけるゼロ和ゲームの文脈で、一貫性のある多様性定義の欠如に取り組むこと。
  • 行動的多様性と応答的多様性を、マルコフゲームに適用可能な理論的根拠に基づく単一の測定値に統合すること。
  • マルチエージェント学習における多様性促進目的を通じて、集団の効果性を向上させ、脆弱性を低減すること。
  • 集団効果性を、脆弱性の代替としてより公平な指標として導入すること。
  • 本手法を行列ゲーム、非推移的混合モデル、Google Research Football環境において検証すること。

提案手法

  • 行動的多様性(BD)を、状態行動空間におけるポリシーの占有測度間の f-発散として定義する。
  • 応答的多様性(RD)を、ポリシーの応答ベクトルとゲームスケープの凸包との幾何的距離として定義し、実用的最適化のための下界を設ける。
  • BD と RD を組み合わせた統一的多様性目的関数を定式化し、学習可能な重み λ₁ と λ₂ を用いて両成分のバランスを調整する。
  • メタゲームにおけるナッシュ均衡を用いて、相手混合ポリシーを計算し、最良応答の計算を支援する。
  • 強化学習ベースの最適化または簡略化された勾配ベースの最適化(例:定理1および定理2を介して)を用い、多様性が高く、効果的な応答へとポリシーを更新する。
  • 2種類のアルゴリズムバージョンを実装する:行列ゲーム用(アルゴリズム3)と微分ゲーム用(アルゴリズム4)で、それぞれBDとRDのための別個の最適化目的関数を有する。

実験結果

リサーチクエスチョン

  • RQ1ゼロ和ゲームにおけるマルチエージェントオープンエンドドラーニングにおいて、行動的多様性と応答的多様性を正式に統一する単一の測定値として定式化することは可能か?
  • RQ2BD と RD を同時に最適化することで、非推移的ゲームにおける集団の脆弱性と効果性にどのような影響を与えるか?
  • RQ3本手法で提案された集団効果性指標は、脆弱性と比較して、ポリシー集団の強さを評価する上で、どのように優れているか?
  • RQ4Google Research Football のような複雑な環境において、統一的多様性測定値は性能向上にどの程度寄与するか?
  • RQ5ハイパーパrameter λ₁ と λ₂ によって制御される行動的多様性と応答的多様性の最適なバランスは何か?

主な発見

  • 行列ゲームにおいて、本手法はベースラインと比較して最も低い脆弱性と最も高い集団効果性を達成した。PSRO に BD と RD を組み合わせたバージョンは、RD のみを用いた PSRO よりも優れた性能を示した。
  • 非推移的混合モデルでは、単一の多様性タイプに依存する手法と比較して、統一的多様性アプローチがより効果的かつ頑健なポリシー集団を生成した。
  • Google Research Football では、異なるスキルレベルの相手に対して最大のゴール差を記録し、優れた一般化能力と適応性を示した。
  • アブレーションスタディの結果、λ₂(応答的多様性重み)を 0.5 に設定した場合が最良の性能を示し、両多様性タイプがバランス良く寄与していることが示された。
  • 集団効果性指標は、特に非推移的状況において、脆弱性よりも公平でより情報量が多い評価指標であることが実証された。
  • 統一枠組みは、ポリシー行動と応答ダイナミクスの両方を効果的に捉えており、複数の環境で理論的および実証的妥当性が確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。