[论文解读] Stochastic Shortest Path Games and Q-Learning
本文建立了在有限状态与紧致控制的两玩家零和随机最短路径(SSP)博弈中,Q-learning 的收敛性。通过引入对称模型条件,确保贝尔曼方程解的唯一性与Q-learning 迭代值的有界性,证明了在温和且可验证的假设下,Q-learning 几乎必然收敛,将无模型学习扩展至未折扣的总成本随机博弈。
We consider a class of two-player zero-sum stochastic games with finite state and compact control spaces, which we call stochastic shortest path (SSP) games. They are undiscounted total cost stochastic dynamic games that have a cost-free termination state. Exploiting the close connection of these games to single-player SSP problems, we introduce novel model conditions under which we show that the SSP games have strong optimality properties, including the existence of a unique solution to the dynamic programming equation, the existence of optimal stationary policies, and the convergence of value and policy iteration. We then focus on finite state and control SSP games and the classical Q-learning algorithm for computing the value function. Q-learning is a model-free, asynchronous stochastic iterative algorithm. By the theory of stochastic approximation involving monotone nonexpansive mappings, it is known to converge when its associated dynamic programming equation has a unique solution and its iterates are bounded with probability one. For the SSP case, as the main result of this paper, we prove the boundedness of the Q-learning iterates under our proposed model conditions, thereby establishing completely the convergence of Q-learning for a broad class of total cost finite-space stochastic games.
研究动机与目标
- 建立两玩家零和SSP博弈的强最优性性质——包括解的唯一性、最优平稳策略的存在性,以及价值/策略迭代的收敛性。
- 识别在无模型、异步设置下,有限状态、紧致控制SSP博弈中Q-learning收敛的模型条件。
- 将Q-learning的理论基础从单智能体MDP扩展至具有总成本准则的两玩家随机博弈。
- 证明Q-learning迭代值几乎必然有界,这是未折扣随机博弈中收敛性所缺失的关键环节。
- 提出一种对称的模型条件表述,推广文献中先前的非对称假设。
提出的方法
- 提出一种新颖的对称模型条件表述(假设2.3),确保SSP博弈中动态规划方程解的唯一性。
- 分析SSP博弈与单智能体SSP问题之间的联系,以利用MDP理论中已知的收敛结果。
- 利用涉及单调非扩张映射的随机逼近理论,分析Q-learning的收敛性。
- 通过特定权函数ξ定义加权上确界范数,证明贝尔曼算子Fν̄的压缩性质。
- 通过将Q-learning迭代值与邻近SSP问题中的总成本过程关联,并利用耦合论证证明其下界与上界有界,从而建立Q-learning迭代值的有界性。
- 应用Tsitsiklis(1994)关于带压缩映射的异步随机逼近的结果,得出几乎必然有界性与收敛性的结论。
实验结果
研究问题
- RQ1两玩家零和SSP博弈的贝尔曼方程在何种条件下具有唯一解?
- RQ2在无模型、异步学习设置下,Q-learning是否能在有限状态、紧致控制的SSP博弈中收敛?
- RQ3何种模型条件可确保Q-learning迭代值在这些博弈中几乎必然有界?
- RQ4如何在未折扣总成本随机博弈中建立Q-learning的收敛性,而此前结果存在局限?
- RQ5何种对称条件可推广SSP博弈理论中先前的非对称假设?
主要发现
- 所提出的模型条件(假设2.3)确保SSP博弈的动态规划方程存在唯一解。
- 在这些条件下,双方玩家均存在最优平稳策略,且价值迭代与策略迭代均收敛。
- 在所提条件下,Q-learning迭代值几乎必然有界,这是收敛性的关键要求。
- Q-learning的收敛性通过随机逼近理论得以建立,其核心依赖于加权范数下贝尔曼算子的压缩性质。
- 有界性证明依赖于将Q-learning迭代值与邻近SSP问题中的总成本过程关联,并引入一种新颖的权函数ξ。
- 该结果将Q-learning收敛性扩展至一大类具有总成本准则的两玩家零和随机博弈,在温和且可验证的假设下成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。