[论文解读] Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments
本文在遍历性和正性条件下,建立了非马尔可夫性、随机控制环境下Q-learning的通用收敛定理。证明了Q-learning在连续状态空间MDP、滤波稳定的量化信念MDP、有限窗口POMDP以及多智能体系统中收敛至近优解,实现了在部分或无模型知识条件下的学习。
As a primary contribution, we present a convergence theorem for stochastic iterations, and in particular, Q-learning iterates, under a general, possibly non-Markovian, stochastic environment. Our conditions for convergence involve an ergodicity and a positivity criterion. We provide a precise characterization on the limit of the iterates and conditions on the environment and initializations for convergence. As our second contribution, we discuss the implications and applications of this theorem to a variety of stochastic control problems with non-Markovian environments involving (i) quantized approximations of fully observed Markov Decision Processes (MDPs) with continuous spaces (where quantization break down the Markovian structure), (ii) quantized approximations of belief-MDP reduced partially observable MDPS (POMDPs) with weak Feller continuity and a mild version of filter stability (which requires the knowledge of the model by the controller), (iii) finite window approximations of POMDPs under a uniform controlled filter stability (which does not require the knowledge of the model), and (iv) for multi-agent models where convergence of learning dynamics to a new class of equilibria, subjective Q-learning equilibria, will be studied. In addition to the convergence theorem, some implications of the theorem above are new to the literature and others are interpreted as applications of the convergence theorem. Some open problems are noted.
研究动机与目标
- 在非马尔可夫性和通用信息结构下,建立Q-learning在随机控制中的通用收敛定理。
- 在遍历性和正性条件下刻画Q-learning迭代的极限,即使底层过程并非马尔可夫过程。
- 将Q-learning的适用范围扩展至连续状态空间MDP、部分可观察MDP(POMDP)以及具有有限或局部信息的多智能体系统。
- 引入并分析一类新均衡——主观Q-learning均衡——在去中心化随机控制设置中的性质。
- 识别在量化或有限记忆导致马尔可夫性被破坏的系统中,Q-learning实现近优性的条件。
提出的方法
- 提出在遍历性和正性条件下,适用于非马尔可夫环境的随机迭代(包括Q-learning)的通用收敛定理。
- 将该收敛定理应用于具有连续空间的完全可观测MDP,利用弱连续性和遍历性条件实现近优性能。
- 在弱Feller连续性和温和滤波稳定性条件下,分析POMDP的信念MDP近似,实现基于量化方法的模型知识学习。
- 在统一受控滤波稳定性条件下,考虑POMDP的有限窗口近似,实现无需模型知识的模型自由学习,且初始值任意。
- 通过双时标、满足性导向的学习方案,将Q-learning算法适配至多智能体系统,结合策略更新与值函数估计。
- 采用主观Q-learning均衡框架,使智能体在无全局信息条件下,仅依赖本地探索与值更新,即可收敛至稳定策略组合。

实验结果
研究问题
- RQ1在何种一般条件下,Q-learning能在非马尔可夫性、随机控制环境中收敛?
- RQ2当马尔可夫性因量化而被破坏时,如何在连续状态空间MDP中应用Q-learning?
- RQ3在量化信念表示和滤波稳定性条件下,POMDP中确保收敛至近优策略的条件是什么?
- RQ4在无模型知识条件下,能否通过有限记忆近似实现在POMDP中的Q-learning近优性?
- RQ5在去中心化多智能体系统中,主观Q-learning均衡的性质及其存在条件是什么?
主要发现
- 即使在非马尔可夫环境中,Q-learning仍收敛至由遍历性和正性条件刻画的极限。
- 对于连续状态空间MDP,Q-learning在弱连续性和遍历性条件下可实现近优性。
- 在POMDP的信念MDP近似中,当滤波动态满足弱Feller连续性且滤波稳定性成立时,Q-learning收敛至近优策略。
- 在POMDP的有限窗口近似中,统一受控滤波稳定性条件可实现无需模型知识的模型自由Q-learning收敛。
- 在多智能体系统中,即使仅具有严格本地信息,Q-learning动态在满足性与双时标学习下仍收敛至主观Q-learning均衡。
- 主观Q-learning均衡的存在性仍为开放问题,但通过算法设计与稳定性分析,已建立其收敛条件。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。