Skip to main content
QUICK REVIEW

[Paper Review] On Markov Decision Processes with Borel Spaces and an Average Cost Criterion

Huizhen Yu|arXiv (Cornell University)|Jan 10, 2019
Economic theories and models37 references4 citations
TL;DR

This paper establishes the average cost optimality inequality (ACOI) for Markov decision processes with Borel state and action spaces under general conditions that do not require compactness or continuity. By introducing majorization-type conditions and leveraging Egoroff’s and Lusin’s theorems, it proves ACOI for unbounded and nonnegative cost models, enabling optimality results in discontinuous dynamics and cost functions.

ABSTRACT

We consider average-cost Markov decision processes (MDPs) with Borel state and action spaces and universally measurable policies. For the nonnegative cost model and an unbounded cost model, we introduce a set of conditions under which we prove the average cost optimality inequality (ACOI) via the vanishing discount factor approach. Unlike most existing results on the ACOI, which require compactness/continuity conditions on the MDP, our result does not and can be applied to problems with discontinuous dynamics and one-stage costs. The key idea here is to replace the compactness/continuity conditions used in the prior work by what we call majorization type conditions. In particular, among others, we require that for each state, on selected subsets of actions at that state, the state transition stochastic kernel is majorized by finite measures, and we use this majorization property together with Egoroff's theorem to prove the ACOI. We also consider the minimum pair approach for average-cost MDPs and apply the majorization idea. For the case of a discrete action space and strictly unbounded costs, we prove the existence of a minimum pair that consists of a stationary policy and an invariant probability measure induced by the policy. This result is derived by combining Lusin's theorem with another majorization condition we introduce, and it can be applied to a class of countable action space MDPs in which, with respect to the state variable, the dynamics and one-stage costs are discontinuous.

Motivation & Objective

  • To extend average cost optimality theory to general Borel-space MDPs without requiring compactness or continuity conditions.
  • To resolve measurability and convergence issues in MDPs with discontinuous dynamics and one-stage costs.
  • To establish the average cost optimality inequality (ACOI) using the vanishing discount factor approach under weaker structural assumptions.
  • To prove existence of a minimum pair (stationary policy and invariant measure) for discrete action spaces with unbounded costs.
  • To develop a framework based on majorization conditions that replace traditional continuity and compactness requirements.

Proposed method

  • Introduces majorization-type conditions where the state transition kernel is dominated by finite measures on selected action subsets.
  • Applies Egoroff’s theorem to extract sets of large measure where functions exhibit uniform convergence properties.
  • Uses Lusin’s theorem to construct measurable selections for policies in the minimum pair existence proof.
  • Employs the vanishing discount factor approach to derive the ACOI by taking limits of discounted cost optimality equations.
  • Combines majorization with Lyapunov-type conditions to ensure integrability and stability in unbounded cost models.
  • Uses measurable selection theorems (e.g., [2, Prop. 7.50]) to construct universally measurable policies from suboptimal actions.

Experimental results

Research questions

  • RQ1Can the average cost optimality inequality (ACOI) be established in Borel-space MDPs without compactness or continuity assumptions?
  • RQ2What alternative conditions can replace continuity and compactness to ensure convergence and optimality in average cost MDPs?
  • RQ3Under what conditions does a minimum pair (stationary policy and invariant measure) exist for MDPs with unbounded costs?
  • RQ4How can Egoroff’s and Lusin’s theorems be leveraged to handle discontinuities in dynamics and cost functions?
  • RQ5Can the vanishing discount factor approach be adapted to prove ACOI under majorization-type conditions?

Key findings

  • The ACOI is established for nonnegative cost models and unbounded cost models with a Lyapunov-type condition, without requiring continuity or compactness.
  • The proof relies on majorization conditions where the transition kernel is dominated by finite measures on action subsets, enabling application of Egoroff’s theorem.
  • For discrete action spaces with strictly unbounded costs, a minimum pair consisting of a stationary policy and an invariant measure exists, proven via Lusin’s theorem and a new majorization condition.
  • The constructed nonrandomized Markov policy is average-cost optimal, and stationary policies are shown to be ε-optimal under the ACOI.
  • The value function and cost-to-go functions are shown to be finite and measurable under the introduced majorization and Lyapunov conditions.
  • The contraction property of the discounted cost operator is established in a weighted norm, ensuring convergence of the value iteration process.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.