Skip to main content
QUICK REVIEW

[论文解读] On Markov Decision Processes with Borel Spaces and an Average Cost Criterion

Huizhen Yu|arXiv (Cornell University)|Jan 10, 2019
Economic theories and models参考文献 37被引用 4
一句话总结

本文在不需紧致性或连续性条件的一般条件下,为具有Borel状态空间和动作空间的马尔可夫决策过程建立了平均成本最优不等式(ACOI)。通过引入类似控制的条件,并利用Egoroff定理和Lusin定理,证明了无界且非负代价模型下的ACOI,从而实现了在不连续动态和代价函数下的最优性结果。

ABSTRACT

We consider average-cost Markov decision processes (MDPs) with Borel state and action spaces and universally measurable policies. For the nonnegative cost model and an unbounded cost model, we introduce a set of conditions under which we prove the average cost optimality inequality (ACOI) via the vanishing discount factor approach. Unlike most existing results on the ACOI, which require compactness/continuity conditions on the MDP, our result does not and can be applied to problems with discontinuous dynamics and one-stage costs. The key idea here is to replace the compactness/continuity conditions used in the prior work by what we call majorization type conditions. In particular, among others, we require that for each state, on selected subsets of actions at that state, the state transition stochastic kernel is majorized by finite measures, and we use this majorization property together with Egoroff's theorem to prove the ACOI. We also consider the minimum pair approach for average-cost MDPs and apply the majorization idea. For the case of a discrete action space and strictly unbounded costs, we prove the existence of a minimum pair that consists of a stationary policy and an invariant probability measure induced by the policy. This result is derived by combining Lusin's theorem with another majorization condition we introduce, and it can be applied to a class of countable action space MDPs in which, with respect to the state variable, the dynamics and one-stage costs are discontinuous.

研究动机与目标

  • 将平均成本最优性理论扩展至一般Borel空间MDP,而无需紧致性或连续性条件。
  • 解决具有不连续动态和单阶段代价的MDP中的可测性与收敛性问题。
  • 在较弱的结构假设下,通过消失折扣因子方法建立平均成本最优不等式(ACOI)。
  • 证明在具有无界代价的离散动作空间中,存在最小配对(平稳策略与不变测度)。
  • 基于控制条件构建一个框架,以替代传统的连续性与紧致性要求。

提出的方法

  • 引入类似控制的条件,其中状态转移核在选定的动作子集上被有限测度控制。
  • 应用Egoroff定理,提取具有大测度的集合,使函数在这些集合上表现出一致收敛性质。
  • 使用Lusin定理,在最小配对存在性证明中构造可测选择。
  • 采用消失折扣因子方法,通过取折扣代价最优方程的极限推导ACOI。
  • 将控制条件与Lyapunov型条件结合,确保无界代价模型中的可积性与稳定性。
  • 使用可测选择定理(例如,[2, Prop. 7.50])从次优动作构造普遍可测策略。

实验结果

研究问题

  • RQ1是否可以在不依赖紧致性或连续性假设的Borel空间MDP中建立平均成本最优不等式(ACOI)?
  • RQ2何种替代条件可取代连续性与紧致性,以确保平均成本MDP中的收敛性与最优性?
  • RQ3在具有无界代价的MDP中,最小配对(平稳策略与不变测度)在何种条件下存在?
  • RQ4如何利用Egoroff定理与Lusin定理处理动态与代价函数中的不连续性?
  • RQ5消失折扣因子方法是否可被调整以在控制型条件下证明ACOI?

主要发现

  • 在非负代价模型及满足Lyapunov型条件的无界代价模型中,ACOI被建立,且无需连续性或紧致性条件。
  • 证明依赖于控制条件,其中转移核在动作子集上被有限测度控制,从而可应用Egoroff定理。
  • 对于具有严格无界代价的离散动作空间,通过Lusin定理与新引入的控制条件,证明了最小配对(包括平稳策略与不变测度)的存在性。
  • 所构造的非随机马尔可夫策略为平均成本最优策略,且在ACOI下,平稳策略被证明为ε-最优。
  • 在引入的控制条件与Lyapunov条件下,值函数与代价-至-目标函数被证明为有限且可测。
  • 在加权范数下,证明了折扣代价算子的压缩性质,从而确保值迭代过程的收敛性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。