Skip to main content
QUICK REVIEW

[論文レビュー] From Motor Control to Team Play in Simulated Humanoid Football

Siqi Liu, Guy Lever|arXiv (Cornell University)|May 25, 2021
Sports Analytics and Performance被引用数 12
ひとこと要約

本論文では、模倣学習、マルチエージェント強化学習、およびパopulationベースのトレーニングを統合することで、物理的に現実的な環境で人間のようないじりのないフットボールをプレーするシミュレーテッドヒューマノイドエージェントを訓練する階層的強化学習フレームワークを提示する。この手法により、ミリ秒単位の低レベルのモーターコントロールから数秒単位の高レベルのチーム連携までを学習可能となり、発生的で人間らしいチーム戦術と、転送可能な行動表現を有する協調的で現実的なフットボールプレーが実現される。

ABSTRACT

Intelligent behaviour in the physical world exhibits structure at multiple spatial and temporal scales. Although movements are ultimately executed at the level of instantaneous muscle tensions or joint torques, they must be selected to serve goals defined on much longer timescales, and in terms of relations that extend far beyond the body itself, ultimately involving coordination with other agents. Recent research in artificial intelligence has shown the promise of learning-based approaches to the respective problems of complex movement, longer-term planning and multi-agent coordination. However, there is limited research aimed at their integration. We study this problem by training teams of physically simulated humanoid avatars to play football in a realistic virtual environment. We develop a method that combines imitation learning, single- and multi-agent reinforcement learning and population-based training, and makes use of transferable representations of behaviour for decision making at different levels of abstraction. In a sequence of stages, players first learn to control a fully articulated body to perform realistic, human-like movements such as running and turning; they then acquire mid-level football skills such as dribbling and shooting; finally, they develop awareness of others and play as a team, bridging the gap between low-level motor control at a timescale of milliseconds, and coordinated goal-directed behaviour as a team at the timescale of tens of seconds. We investigate the emergence of behaviours at different levels of abstraction, as well as the representations that underlie these behaviours using several analysis techniques, including statistics from real-world sports analytics. Our work constitutes a complete demonstration of integrated decision-making at multiple scales in a physically embodied multi-agent setting. See project video at https://youtu.be/KHMwq9pv7mg.

研究の動機と目的

  • 身体的でマルチエージェントシステムにおける低レベルのモーターコントロールと高レベルのチーム連携のギャップを埋めること。
  • 複雑でマルチスケールの行動を実現するため、模倣学習、単一エージェントおよびマルチエージェント強化学習、およびパopulationベースのトレーニングの統合を調査すること。
  • 物理的に現実的で人間らしい動きと、発生的チーム戦術を有するシミュレーテッドフットボール環境で実現すること。
  • 実世界のスポーツアナリティクス手法を用いて、複数の時間的・空間的スケールにおける行動の発生を分析すること。
  • 挑戦的でマルチエージェントかつ物理ベースの環境において、協調的で長時間にわたる行動のエンドツーエンド学習を実証すること。

提案手法

  • まず、モーションキャプチャデータからの模倣学習により、基本的な歩行を段階的トレーニングで学習する。
  • 単一エージェント強化学習を適用し、ドリブルやシュートといった中レベルのスキルを発展させる。
  • 自己対戦を用いたマルチエージェント強化学習を採用し、チーム連携と戦術的認識を発展させる。
  • 異なるスキルレベルやエージェント間での知識転送を可能にするために、転送可能な行動表現を導入する。
  • ハイパーパrameter最適化と多様なチーム構成における一般化の向上を図るため、パopulationベースのトレーニングを活用する。
  • 物理的相互作用(身体的接触を含む)を含む複雑な相互作用をサポートする、物理ベースのシミュレーション環境を採用する。

実験結果

リサーチクエスチョン

  • RQ1マルチエージェントで物理的にシミュレートされた環境において、低レベルのモーターコントロールを高レベルのチーム連携と効果的に統合する方法は何か?
  • RQ2転送可能な行動表現は、複数の抽象化レベルにわたる階層的学習をどのように可能にするか?
  • RQ3模倣学習と強化学習を統合することで、自然で人間らしい動きと効果的なチームプレイを両立させられるか?
  • RQ4シミュレーテッドフットボールにおける発生的チーム戦術は、実世界のスポーツアナリティクスのパターンとどのように比較できるか?
  • RQ5エンドツーエンド学習手法は、複雑でマルチエージェントな環境において、協調的で長時間にわたる行動の発生をどの程度サポートできるか?

主な発見

  • 模倣学習により、走る、方向転換する、ドリブルするといった現実的で人間らしい動きが効果的に学習された。
  • 単一エージェント強化学習を用いたシミュレーション環境で、シュートやパスといった中レベルのフットボールスキルが自然に発生した。
  • 自己対戦を用いたマルチエージェント強化学習により、チーム連携と戦術的認識が発展し、効果的な2対2フットボールプレイが実現された。
  • 階層的スキル学習と転送可能な表現の統合により、ミリ秒から数秒のスケールまで、複数の時間スケールにわたる効率的な学習が可能になった。
  • 実世界のスポーツアナリティクスを用いた分析から、発生的チーム行動がプロフェッショナルフットボールで観察されるパターンと密接に類似していることが明らかになった。
  • 本フレームワークは、カリキュラムベースのトレーニングを用いたエンドツーエンド学習が、マルチエージェントシステムにおける複雑で協調的かつ物理的に根拠のある行動をサポートできることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。