[Paper Review] On the computational complexity of solving stochastic mean-payoff games
This paper establishes polynomial-time equivalence between solving Condon games (simple stochastic games), stochastic mean-payoff games with unary-reward/probability representations, and the same games with binary representations. The authors prove that the computational complexity of solving these three classes of games is identical, implying that adding rewards and using binary encoding does not increase the problem's difficulty beyond that of standard simple stochastic games.
We consider some well-known families of two-player, zero-sum, perfect information games that can be viewed as special cases of Shapley's stochastic games. We show that the following tasks are polynomial time equivalent: - Solving simple stochastic games. - Solving stochastic mean-payoff games with rewards and probabilities given in unary. - Solving stochastic mean-payoff games with rewards and probabilities given in binary.
Motivation & Objective
- To determine whether solving stochastic mean-payoff games (undiscounted Gillette games) is computationally equivalent to solving simple stochastic games (Condon games).
- To investigate whether the inclusion of rewards during gameplay increases the computational complexity of solving stochastic games.
- To examine the impact of input representation—specifically unary versus binary encoding of rewards and probabilities—on the complexity of solving these games.
- To explore whether further generalizations of Gillette games, such as Filar games with one-player-dependent transitions, remain polynomial-time equivalent to Condon games.
Proposed method
- Construct a reduction from discounted Gillette games to Condon games by introducing auxiliary vertices and probabilistic gadgets to simulate discounting via terminal absorption probabilities.
- Affinely scale all rewards in the original game to the [0,1] interval to preserve strategy equivalence without altering optimal behavior.
- For each (state, action) pair in the original game, create a new random vertex in the Condon game that routes play based on the action’s reward and transition probabilities, with absorption into the terminal vertex with probability proportional to the discounted reward.
- Demonstrate strategic equivalence by showing that the probability of reaching the terminal state in the Condon game equals the discounted expected reward in the original Gillette game.
- Construct a reverse reduction from Condon games to undiscounted Gillette games by mapping terminal absorption to a reward of 1 and all other transitions to reward 0, preserving the probability of reaching the terminal state as the limiting average payoff.
- Prove that pure, positional strategies in the original game correspond one-to-one with those in the transformed game, and that optimal strategies are preserved under both reductions.
Experimental results
Research questions
- RQ1Are solving Condon games and solving stochastic mean-payoff games (undiscounted Gillette games) computationally equivalent?
- RQ2Does the inclusion of rewards during gameplay in stochastic games increase the computational complexity beyond that of simple stochastic games?
- RQ3Is the computational complexity of solving stochastic mean-payoff games dependent on the input representation (unary vs. binary encoding of rewards and probabilities)?
- RQ4Can further generalizations of Gillette games—such as Filar games with one-player-dependent transitions—be reduced to Condon games in polynomial time?
Key findings
- Solving Condon games is polynomial-time equivalent to solving stochastic mean-payoff games with rewards and probabilities represented in unary.
- Solving stochastic mean-payoff games with rewards and probabilities in binary is also polynomial-time equivalent to solving Condon games.
- The presence of rewards during gameplay does not increase the computational complexity of solving stochastic games beyond that of simple stochastic games.
- The reduction from discounted Gillette games to Condon games preserves optimal strategies and is computable in polynomial time, with the absorption probability in the Condon game equal to the discounted reward in the original game.
- The reduction from Condon games to undiscounted Gillette games preserves the probability of reaching the terminal state as the limiting average payoff, ensuring strategic equivalence.
- The results imply that the computational complexity of solving these game families is fundamentally the same, regardless of whether rewards are included or whether inputs are encoded in unary or binary.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.