[论文解读] Policy iteration algorithm for zero-sum multichain stochastic games with mean payoff and perfect information
本文提出了一种针对具有平均收益和完美信息的零和多链随机博弈的策略迭代算法,利用非线性谱投影以确保在退化迭代下仍能收敛。该方法通过词典序单调性保证终止,并成功求解大规模实例,包括Richman博弈和追逃博弈,且验证了平均收益向量与相对值向量。
We consider zero-sum stochastic games with finite state and action spaces, perfect information, mean payoff criteria, without any irreducibility assumption on the Markov chains associated to strategies (multichain games). The value of such a game can be characterized by a system of nonlinear equations, involving the mean payoff vector and an auxiliary vector (relative value or bias). We develop here a policy iteration algorithm for zero-sum stochastic games with mean payoff, following an idea of two of the authors (Cochet-Terrasson and Gaubert, C. R. Math. Acad. Sci. Paris, 2006). The algorithm relies on a notion of nonlinear spectral projection (Akian and Gaubert, Nonlinear Analysis TMA, 2003), which is analogous to the notion of reduction of super-harmonic functions in linear potential theory. To avoid cycling, at each degenerate iteration (in which the mean payoff vector is not improved), the new relative value is obtained by reducing the earlier one. We show that the sequence of values and relative values satisfies a lexicographical monotonicity property, which implies that the algorithm does terminate. We illustrate the algorithm by a mean-payoff version of Richman games (stochastic tug-of-war or discrete infinity Laplacian type equation), in which degenerate iterations are frequent. We report numerical experiments on large scale instances, arising from the latter games, as well as from monotone discretizations of a mean-payoff pursuit-evasion deterministic differential game.
研究动机与目标
- 开发一种针对具有平均收益和完美信息的零和多链随机博弈的策略迭代算法,且不对马尔可夫链做不可约性假设。
- 在平均收益向量未改善的退化迭代存在的情况下,确保算法的收敛性。
- 通过涉及平均收益向量和相对值向量(偏差)的非线性方程组刻画博弈的价值。
- 将该算法应用于实际问题,如Richman博弈和追逃微分博弈的单调离散化。
- 展示该算法在大规模数值实例上的效率与鲁棒性。
提出的方法
- 在退化迭代期间,该算法使用一种类似于线性势论中超调和函数约化的非线性谱投影来更新相对值向量。
- 在每次迭代中,算法计算当前策略的平均收益和相对值,当平均收益未改善时,通过一个约减步骤来更新后者。
- 平均收益和相对值向量序列满足词典序单调性,确保不会出现循环且能有限终止。
- 动态规划算子是多面体的、保序的且可加齐次的,使得能够使用不变半直线刻画方法来表征价值。
- 该算法通过双人策略迭代循环实现,交替优化第一方和第二方的策略。
- 数值实验采用连续追逃博弈的有限差分离散化,状态空间在网格上离散化,动作受限以避免边界违规。
实验结果
研究问题
- RQ1能否为具有平均收益的零和多链随机博弈设计一种策略迭代算法,即使在频繁出现退化迭代的情况下也能避免循环?
- RQ2非线性谱投影概念如何被调整以在无不可约性假设下确保收敛?
- RQ3相对值向量在刻画多链博弈中平均收益时起什么作用?
- RQ4该算法在大规模实例(如离散化追逃博弈或Richman博弈)上的表现如何?
- RQ5在速度不对称的博弈(如老鼠与猫博弈)中,相对值向量与最优策略之间有何关系?
主要发现
- 由于价值与相对值序列的词典序单调性,该算法即使在存在退化迭代的情况下也能在有限时间内终止。
- 当猫的速度为0.999时,老鼠与猫博弈中,若x在捕获球外部,则平均收益为η(x) = 0.492;若x在内部,则η(x) = 0,表明老鼠可维持距离。
- 当猫的速度为1.001时,平均收益接近于零,表明猫最终会捕获老鼠,且相对值向量显示出捕获时间的结构。
- 当速度相等(bar{b} = 1)时,相对值近似为零,平均收益为η(x) ≈ ||x||²₂,表明距离保持恒定。
- 该算法成功求解了257×257网格的大规模实例,残差低于1e-14,CPU时间低于1000秒。
- 在bar{b} = 0.999和bar{b} = 1.001时,仅在最后几步出现强退化迭代,表明收敛具有稳定性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。