Skip to main content
QUICK REVIEW

[Paper Review] Enhancements for Real-Time Monte-Carlo Tree Search in General Video Game Playing

Dennis J. N. J. Soemers, Chiara F. Sironi|Data Archiving and Networked Services (DANS)|Jul 3, 2024
Artificial Intelligence in Games10 citations
TL;DR

The paper introduces eight enhancements to open-loop MCTS for GVGP and shows that, individually and combined, they significantly improve win rates across sixty GVGP games, approaching competitive levels from 2015 GVG-AI competition.

ABSTRACT

General Video Game Playing (GVGP) is a field of Artificial Intelligence where agents play a variety of real-time video games that are unknown in advance. This limits the use of domain-specific heuristics. Monte-Carlo Tree Search (MCTS) is a search technique for game playing that does not rely on domain-specific knowledge. This paper discusses eight enhancements for MCTS in GVGP; Progressive History, N-Gram Selection Technique, Tree Reuse, Breadth-First Tree Initialization, Loss Avoidance, Novelty-Based Pruning, Knowledge-Based Evaluations, and Deterministic Game Detection. Some of these are known from existing literature, and are either extended or introduced in the context of GVGP, and some are novel enhancements for MCTS. Most enhancements are shown to provide statistically significant increases in win percentages when applied individually. When combined, they increase the average win percentage over sixty different games from 31.0% to 48.4% in comparison to a vanilla MCTS implementation, approaching a level that is competitive with the best agents of the GVG-AI competition in 2015.

Motivation & Objective

  • Motivate and improve general video game playing agents that must operate without domain-specific heuristics.
  • Evaluate how enhancements to open-loop MCTS affect performance in the GVGP setting.
  • Identify which enhancements yield statistically significant improvements individually and in combination.

Proposed method

  • Describe eight enhancements: Progressive History, N-Gram Selection Technique, Tree Reuse, Breadth-First Tree Initialization, Loss Avoidance, Novelty-Based Pruning, Knowledge-Based Evaluations, and Deterministic Game Detection.
  • Use open-loop MCTS within the GVG-AI framework, with a basic rollout evaluation X(s_T) to backpropagate outcomes.
  • Experimentally compare baseline MCTS and enhanced variants across sixty GVGP games with multiple setups and 95% confidence intervals.
  • Investigate deterministic vs nondeterministic game handling via Deterministic Game Detection and adjust tree reuse and pruning accordingly.

Experimental results

Research questions

  • RQ1Do the eight proposed enhancements improve win rates in GVGP when applied individually?
  • RQ2Do combinations of enhancements yield additive or synergistic improvements over vanilla MCTS in GVGP?
  • RQ3How do enhancements impact the rate of losses and game termination times in GVGP?
  • RQ4How does Deterministic Game Detection influence MCTS behavior and performance in deterministic vs nondeterministic GVGP games.

Key findings

  • Combined enhancements increase average win percentage from 31.0% (vanilla MCTS) to 48.4% across 60 games.
  • Individual enhancements often yield statistically significant win-rate gains over vanilla MCTS.
  • Breadth-First Tree Initialization with Safety Prepruning reduces premature losses and increases robustness in some sets.
  • Knowledge-Based Evaluations, Loss Avoidance, and Novelty-Based Pruning each contribute notable gains, with KBE often providing the largest individual improvement.
  • Deterministic Game Detection informs mixmax-style adjustments and selective pruning for deterministic games.
  • Tree Reuse with appropriate decay (gamma) can improve win rates for certain configurations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.