Skip to main content
QUICK REVIEW

[Paper Review] Fictitious Play for Mean Field Games: Continuous Time Analysis and Applications

Sarah Perrin, Julien Pérolat|arXiv (Cornell University)|Jul 5, 2020
Game Theory and Applications19 citations
TL;DR

This paper introduces a continuous-time Fictitious Play algorithm for Mean Field Games (MFGs) with finite horizon and discounted settings, including common noise. It proves convergence at rate $O(1/t)$ in exploitability, establishes exploitability as a key metric for Nash equilibrium approximation, and provides the first convergence guarantees for learning in MFGs with common noise, validated in both model-based and model-free settings.

ABSTRACT

In this paper, we deepen the analysis of continuous time Fictitious Play learning algorithm to the consideration of various finite state Mean Field Game settings (finite horizon, $γ$-discounted), allowing in particular for the introduction of an additional common noise. We first present a theoretical convergence analysis of the continuous time Fictitious Play process and prove that the induced exploitability decreases at a rate $O(\frac{1}{t})$. Such analysis emphasizes the use of exploitability as a relevant metric for evaluating the convergence towards a Nash equilibrium in the context of Mean Field Games. These theoretical contributions are supported by numerical experiments provided in either model-based or model-free settings. We provide hereby for the first time converging learning dynamics for Mean Field Games in the presence of common noise.

Motivation & Objective

  • To extend Fictitious Play to continuous-time Mean Field Games with finite horizon and $γ$-discounted settings, including common noise.
  • To establish exploitability as a meaningful metric for evaluating convergence to Nash equilibrium in MFGs.
  • To provide theoretical convergence guarantees for Fictitious Play in MFGs, with a rate of $O(1/t)$, matching results in zero-sum games.
  • To demonstrate the algorithm’s performance in both model-based and model-free environments, including with common noise.
  • To bridge the gap between scalable learning algorithms and theoretical convergence in large-population stochastic games.

Proposed method

  • The paper formulates a continuous-time Fictitious Play process where each agent updates its policy based on the empirical distribution of the population’s state over time.
  • It introduces a continuous-time dynamical system that models the evolution of the representative agent’s policy and the population’s distribution, using backward and forward Kolmogorov equations.
  • The analysis leverages tools from continuous-time learning dynamics, including Lyapunov functions and exploitability-based convergence criteria.
  • It defines exploitability as the maximum gain an agent could achieve by deviating from its current policy, given the population distribution.
  • The method incorporates common noise via a joint diffusion process affecting all agents, modeled through a common Brownian motion $W^0_t$.
  • Theoretical convergence is proven using Riccati equations and ODE analysis in linear-quadratic MFGs, serving as a benchmark for empirical validation.

Experimental results

Research questions

  • RQ1Can Fictitious Play in continuous time achieve $O(1/t)$ convergence rate in finite-horizon and $γ$-discounted MFGs?
  • RQ2How can exploitability be generalized and used as a reliable metric for convergence to Nash equilibrium in MFGs?
  • RQ3Is it possible to derive convergence guarantees for Fictitious Play in MFGs with common noise, which are prevalent in real-world applications?
  • RQ4Does the Fictitious Play process converge to a Nash equilibrium in both model-based and model-free settings under the proposed framework?
  • RQ5Can exploitability serve as a practical and theoretically grounded stopping criterion for MFG learning algorithms?

Key findings

  • The Fictitious Play process converges to a Nash equilibrium at a rate of $O(1/t)$ in exploitability for finite-horizon and $γ$-discounted MFGs.
  • Exploitability is established as a meaningful and effective metric for evaluating convergence quality in MFGs, outperforming population distribution-based metrics.
  • This is the first work to prove convergence of a learning algorithm in MFGs with common noise, a setting previously lacking theoretical guarantees.
  • Numerical experiments confirm convergence in both model-based and model-free settings, validating the theoretical findings.
  • The linear-quadratic MFG benchmark with explicit Riccati solutions confirms the accuracy of the learned policies and value functions.
  • Theoretical analysis shows that small exploitability implies small deviation in value functions across time, supporting exploitability as a proxy for equilibrium quality.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.