[Paper Review] Constant payoff in zero-sum stochastic games
This paper proves that in zero-sum stochastic games with patient players, the expected total payoff remains constant regardless of the game state, confirming a conjecture by Sorin, Venel, and Vigeral (2010). The result is established using semi-algebraic methods, rare transition Markov chains, and variational inequalities for value functions.
In a zero-sum stochastic game, at each stage, two adversary players take decisions and receive a stage payoff determined by them and by a random variable representing the state of nature. The total payoff is the discounted sum of the stage payoffs. Assume that the players are very patient and use optimal strategies. We then prove that, at any point in the game, players get essentially the same expected payoff: the payoff is constant. This solves a conjecture by Sorin, Venel and Vigeral (2010). The proof relies on the semi-algebraic approach for discounted stochastic games introduced by Bewley and Kohlberg (1976), on the theory of Markov chains with rare transitions, initiated by Friedlin and Wentzell (1984), and on some variational inequalities for value functions inspired by the recent work of Davini, Fathi, Iturriaga and Zavidovique (2016)
Motivation & Objective
- To resolve the conjecture that expected payoffs remain constant in zero-sum stochastic games when players are very patient.
- To analyze the long-run behavior of discounted stochastic games under optimal play.
- To establish that the value function is invariant across states under the limit of patient behavior.
- To unify techniques from semi-algebraic game theory, rare transition Markov chains, and variational analysis of value functions.
Proposed method
- Adapts the semi-algebraic approach of Bewley and Kohlberg (1976) to analyze the asymptotic structure of value functions in discounted stochastic games.
- Applies the theory of Markov chains with rare transitions by Friedlin and Wentzell (1984) to model slow state changes under patient strategies.
- Uses variational inequalities inspired by Davini et al. (2016) to characterize the limiting behavior of value functions.
- Establishes that the value function converges to a constant across all states in the limit of vanishing discounting.
- Combines these tools to show that optimal strategies yield identical expected payoffs regardless of the current state.
- Demonstrates that the game’s value is independent of the initial state under optimal play in the long-run limit.
Experimental results
Research questions
- RQ1Does the expected total payoff remain constant across all states in a zero-sum stochastic game when players are very patient?
- RQ2Can the conjecture of constant payoff under patient behavior be formally proven using existing analytical tools?
- RQ3How do rare transitions in the state process affect the long-run value structure of stochastic games?
- RQ4To what extent can variational inequalities for value functions describe the limiting behavior of discounted stochastic games?
- RQ5Is the value function invariant across states in the limit of vanishing discounting?
Key findings
- The expected total payoff is constant across all states when players are very patient and play optimally.
- The value function converges to a constant in the limit of vanishing discounting, regardless of the initial state.
- The result confirms the conjecture by Sorin, Venel, and Vigeral (2010) on constant payoffs in zero-sum stochastic games.
- The semi-algebraic framework enables precise characterization of the limiting value structure.
- The integration of rare transition theory and variational inequalities provides a new analytical pathway for long-run game behavior.
- The payoff constancy arises from the interplay of optimal strategies and the asymptotic stability of the state process.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.