Skip to main content
QUICK REVIEW

[論文レビュー] Enhancements for Real-Time Monte-Carlo Tree Search in General Video Game Playing

Dennis J. N. J. Soemers, Chiara F. Sironi|Data Archiving and Networked Services (DANS)|Jul 3, 2024
Artificial Intelligence in Games被引用数 10
ひとこと要約

本論文は、GVGP のためのオープンループ MCTS に対する8つの改善を提案し、それらを個別におよび組み合わせて適用した場合、60の GVGP ゲーム全体で勝率を有意に向上させ、2015年の GVG-AI 大会の競争レベルに近づくことを示している。

ABSTRACT

General Video Game Playing (GVGP) is a field of Artificial Intelligence where agents play a variety of real-time video games that are unknown in advance. This limits the use of domain-specific heuristics. Monte-Carlo Tree Search (MCTS) is a search technique for game playing that does not rely on domain-specific knowledge. This paper discusses eight enhancements for MCTS in GVGP; Progressive History, N-Gram Selection Technique, Tree Reuse, Breadth-First Tree Initialization, Loss Avoidance, Novelty-Based Pruning, Knowledge-Based Evaluations, and Deterministic Game Detection. Some of these are known from existing literature, and are either extended or introduced in the context of GVGP, and some are novel enhancements for MCTS. Most enhancements are shown to provide statistically significant increases in win percentages when applied individually. When combined, they increase the average win percentage over sixty different games from 31.0% to 48.4% in comparison to a vanilla MCTS implementation, approaching a level that is competitive with the best agents of the GVG-AI competition in 2015.

研究の動機と目的

  • ドメイン特有のヒューリスティクスなしで動作する一般的なビデオゲームプレイエージェントを動機づけ、改善する。
  • GVGP設定におけるオープンループ MCTS の改善が性能に与える影響を評価する。
  • 個別および組み合わせで統計的に有意な改善をもたらす改善を特定する。

提案手法

  • 8つの改善を説明する:Progressive History、N-Gram Selection Technique、Tree Reuse、Breadth-First Tree Initialization、Loss Avoidance、Novelty-Based Pruning、Knowledge-Based Evaluations、Deterministic Game Detection。
  • GVG-AI フレームワーク内でオープンループ MCTS を使用し、結果をバックプロパゲーションする基本的な rollout 評価 X(s_T) を用いる。
  • 複数の設定と95%信頼区間を用いて、60 のGVGPゲームを横断してベースライン MCTS と改善版を実験的に比較する。
  • Deterministic Game Detection を用いた決定論的ゲームと非決定論的ゲームの扱いを調査し、それに応じて木の再利用と剪定を調整する。

実験結果

リサーチクエスチョン

  • RQ18つの提案された改善は、個別に適用した場合、GVGP における勝率を改善するか?
  • RQ2改善の組み合わせは、GVGP における通常の MCTS よりも加法的または相乗的な改善をもたらすか?
  • RQ3改善はGVGPにおける敗北率とゲーム終了時間にどのような影響を与えるか?
  • RQ4Deterministic Game Detection は決定論的および非決定論的な GVGP ゲームにおける MCTS の挙動と性能にどのように影響するか。

主な発見

  • 組み合わせた改善は、60ゲームでの平均勝率を、(ベースラインの MCTS) の 31.0% から 48.4% に向上させる。
  • 個別の改善はしばしば vanilla MCTS に対して統計的に有意な勝率向上をもたらす。
  • Breadth-First Tree Initialization は Safety Prepruning を用いることで、早期の敗北を減らし、一部のセットで堅牢性を高める。
  • Knowledge-Based Evaluations、Loss Avoidance、Novelty-Based Pruning はそれぞれ顕著な利得をもたらし、KBE が個別の最大の改善を提供することが多い。
  • Deterministic Game Detection は決定論的ゲームに対して mixmax 風の調整と選択的な剪定を知らせる。
  • Tree Reuse with appropriate decay (gamma) can improve win rates for certain configurations.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。