Skip to main content
QUICK REVIEW

[논문 리뷰] Enhancements for Real-Time Monte-Carlo Tree Search in General Video Game Playing

Dennis J. N. J. Soemers, Chiara F. Sironi|Data Archiving and Networked Services (DANS)|2024. 07. 03.
Artificial Intelligence in Games인용 수 10
한 줄 요약

본 논문은 GVGP를 위한 개방 루프 MCTS에 여덟 가지 개선을 도입하고, 개별적으로 그리고 결합적으로 이들이 60개의 GVGP 게임에서 승률을 크게 향상시키며 2015년 GVG-AI 대회에 근접한 경쟁 수준에 도달함을 보인다.

ABSTRACT

General Video Game Playing (GVGP) is a field of Artificial Intelligence where agents play a variety of real-time video games that are unknown in advance. This limits the use of domain-specific heuristics. Monte-Carlo Tree Search (MCTS) is a search technique for game playing that does not rely on domain-specific knowledge. This paper discusses eight enhancements for MCTS in GVGP; Progressive History, N-Gram Selection Technique, Tree Reuse, Breadth-First Tree Initialization, Loss Avoidance, Novelty-Based Pruning, Knowledge-Based Evaluations, and Deterministic Game Detection. Some of these are known from existing literature, and are either extended or introduced in the context of GVGP, and some are novel enhancements for MCTS. Most enhancements are shown to provide statistically significant increases in win percentages when applied individually. When combined, they increase the average win percentage over sixty different games from 31.0% to 48.4% in comparison to a vanilla MCTS implementation, approaching a level that is competitive with the best agents of the GVG-AI competition in 2015.

연구 동기 및 목표

  • 도메인 특화 휴리스틱 없이 작동해야 하는 일반 비디오 게임 플레이 에이전트를 동기부여하고 향상시키다.
  • GVGP 설정에서 개방 루프 MCTS의 개선이 성능에 미치는 영향을 평가하다.
  • 개별적으로 그리고 조합으로 통계적으로 유의한 개선을 가져오는 개선 요소를 식별하다.

제안 방법

  • 여덟 가지 개선을 설명하다: Progressive History, N-Gram Selection Technique, Tree Reuse, Breadth-First Tree Initialization, Loss Avoidance, Novelty-Based Pruning, Knowledge-Based Evaluations, Deterministic Game Detection.
  • GVG-AI 프레임워크 내에서 개방 루프 MCTS를 사용하고, 결과를 역전파하기 위해 기본 롤아웃 평가 X(s_T)를 적용한다.
  • 여러 구성과 95% 신뢰구간으로 60개의 GVGP 게임에서 기본 MCTS와 향상 변형들을 실험적으로 비교한다.
  • Deterministic Game Detection를 통해 결정적 대 비결정적 게임 처리 방법을 조사하고 그에 따라 트리 재사용과 가지치기를 조정한다.

실험 결과

연구 질문

  • RQ1제안된 여덟 가지 개선이 각각 적용될 때 GVGP에서 승률을 향상시키는가?
  • RQ2개선의 조합이 일반 MCTS에 비해 더해지거나 시너지 효과를 내는 향상을 GVGP에서 가져오는가?
  • RQ3개선들이 GVGP에서 손실 비율과 게임 종료 시간에 어떤 영향을 미치는가?
  • RQ4Deterministic Game Detection이 결정적 게임 대 비결정적 GVGP 게임에서 MCTS 동작 및 성능에 어떤 영향을 미치는가?

주요 결과

  • 결합된 개선은 평균 승률을 31.0%(바닐라 MCTS)에서 48.4%로 올려 60개 게임에서.
  • 개별 개선들도 종종 바닐라 MCTS에 비해 통계적으로 유의한 승률 향상을 보인다.
  • 안전 선처리(Safety Prepruning)를 포함한 너비 우선 트리 초기화는 일부 집합에서 조기 손실을 줄이고 강건성을 높인다.
  • Knowledge-Based Evaluations, Loss Avoidance, Novelty-Based Pruning 각각 주목할 만한 이득을 제공하며, KBE가 종종 가장 큰 개별 개선을 나타낸다.
  • Deterministic Game Detection은 결정적 게임에 대해 mixmax 스타일의 조정과 선별적 가지치를 알리는 역할을 한다.
  • 적절한 감소(gamma)가 있는 Tree Reuse는 특정 구성을 위한 승률을 향상시킬 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.