Skip to main content
QUICK REVIEW

[Paper Review] Convergence of Learning Dynamics in Stackelberg Games

Tanner Fiez, Benjamin J. Chasnov|arXiv (Cornell University)|Jun 4, 2019
Markov Chains and Monte Carlo Methods64 references45 citations
TL;DR

This paper analyzes convergence of gradient-based learning dynamics in Stackelberg games with continuous actions, showing when stable points are Stackelberg equilibria and proposing algorithms with convergence guarantees.

ABSTRACT

This paper investigates the convergence of learning dynamics in Stackelberg games. In the class of games we consider, there is a hierarchical game being played between a leader and a follower with continuous action spaces. We establish a number of connections between the Nash and Stackelberg equilibrium concepts and characterize conditions under which attracting critical points of simultaneous gradient descent are Stackelberg equilibria in zero-sum games. Moreover, we show that the only stable critical points of the Stackelberg gradient dynamics are Stackelberg equilibria in zero-sum games. Using this insight, we develop a gradient-based update for the leader while the follower employs a best response strategy for which each stable critical point is guaranteed to be a Stackelberg equilibrium in zero-sum games. As a result, the learning rule provably converges to a Stackelberg equilibria given an initialization in the region of attraction of a stable critical point. We then consider a follower employing a gradient-play update rule instead of a best response strategy and propose a two-timescale algorithm with similar asymptotic convergence guarantees. For this algorithm, we also provide finite-time high probability bounds for local convergence to a neighborhood of a stable Stackelberg equilibrium in general-sum games. Finally, we present extensive numerical results that validate our theory, provide insights into the optimization landscape of generative adversarial networks, and demonstrate that the learning dynamics we propose can effectively train generative adversarial networks.

Motivation & Objective

  • Motivate and formalize learning dynamics in hierarchical Stackelberg games where a leader and follower interact with continuous action spaces.
  • Characterize connections between Nash and Stackelberg equilibria in zero-sum and general-sum settings.
  • Develop gradient-based learning rules that ensure convergence to Stackelberg equilibria under suitable conditions.
  • Provide analysis for both exact best-response followers and gradient-play followers, with timescale separation between leader and follower.
  • Demonstrate applicability to generative adversarial networks and validate theory with numerical experiments.

Proposed method

  • Define differential Stackelberg equilibrium as a local notion amenable to computation (Definition 4).
  • Derive and analyze leader-follower gradient updates with implicit follower reactions (Eq. (2) and related expressions).
  • Establish that the only stable critical points of Stackelberg gradient dynamics in zero-sum games are Stackelberg equilibria (Proposition 1).
  • Show that stable differential Nash equilibria in zero-sum games are differential Stackelberg equilibria (Proposition 2).
  • Propose a two-timescale algorithm where the follower uses gradient-play; prove almost sure convergence to Stackelberg equilibria in zero-sum games and to stable attractors in general-sum games; provide finite-time high-probability bounds for local convergence.
  • Relate findings to GANs and discuss implications for optimization landscapes in adversarial learning.

Experimental results

Research questions

  • RQ1Under what conditions do attractors of simultaneous gradient play correspond to Stackelberg equilibria in zero-sum and general-sum games?
  • RQ2Can gradient-based Stackelberg learning dynamics converge to Stackelberg equilibria when the follower uses best-response versus gradient-play updates?
  • RQ3What is the relationship between Nash equilibria and Stackelberg equilibria in zero-sum and general-sum settings under the proposed dynamics?
  • RQ4Do the proposed dynamics avoid non-Nash attractors that plague simultaneous gradient descent, and can they guarantee convergence to Stackelberg equilibria in GAN-like scenarios?

Key findings

  • Attracting critical points of the Stackelberg gradient dynamics in zero-sum games are Stackelberg equilibria (Proposition 1).
  • Stable differential Nash equilibria in zero-sum games are differential Stackelberg equilibria (Proposition 2).
  • There exist stable attractors of simultaneous gradient play that are Stackelberg equilibria but not Nash equilibria, with necessary and sufficient conditions identified for when convergence to Stackelberg equilibria occurs (Propositions 3–4).
  • Specializations to GANs under realizable assumptions show conditions where Stackelberg equilibria describe the learning landscape in GANs (Propositions 5–6).
  • A two-timescale algorithm with a gradient-play follower yields almost sure convergence to Stackelberg equilibria in zero-sum games and to stable attractors in general-sum games, plus finite-time high-probability convergence bounds.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.