Skip to main content
QUICK REVIEW

[Paper Review] A New Condition for the Existence of Optimal Stationary Policies in Denumerable State Average Cost Continuous Time Markov Decision Processes with Unbounded Cost and Transition Rates

Ping Cao, Jingui Xie|arXiv (Cornell University)|Apr 22, 2015
Fault Detection and Control Systems7 references3 citations
TL;DR

This paper introduces a new existence condition for optimal stationary policies in denumerable-state continuous-time Markov decision processes (CTMDPs) with unbounded costs and transition rates. By linking the existence of an optimal policy to the finiteness of expected first-passage times and costs to a reference state, the method leverages established queueing system stability results to bypass complex function constructions, significantly simplifying verification and ensuring the average-cost optimality equation holds under mild conditions.

ABSTRACT

This paper presents a new condition for the existence of optimal stationary policies in average-cost continuous-time Markov decision processes with unbounded cost and transition rates, arising from controlled queueing systems. This condition is closely related to the stability of queueing systems. It suggests that the proof of the stability can be exploited to verify the existence of an optimal stationary policy. This new condition is easier to verify than existing conditions. Moreover, several conditions are provided which suffice for the average-cost optimality equality to hold.

Motivation & Objective

  • To address the challenge of verifying optimal stationary policies in continuous-time Markov decision processes with unbounded costs and transition rates.
  • To develop a condition for the existence of average-cost optimal stationary policies that is easier to verify than existing function-based approaches.
  • To connect the existence of optimal policies to the stability of underlying queueing systems, enabling reuse of existing stability analysis.
  • To establish sufficient conditions under which the average-cost optimality equation (ACOE) holds.

Proposed method

  • Proposes a new existence condition based on the finiteness of expected first-passage time and expected cost from any state to a reference state under a given policy.
  • Utilizes the concept of a 'standard policy'—specifically, a policy that is i₀-standard if expected first-passage time and cost to state i₀ are finite.
  • Applies a Lyapunov function approach with exponential weighting to verify the finiteness of expected cost, using inequalities involving transition rates and cost functions.
  • Employs the uniformization method in theory but focuses on unbounded-rate cases where it fails, justifying direct CTMDP analysis.
  • Establishes that if a policy is i₀-standard and satisfies certain drift conditions, the average-cost optimality inequality (ACOI) holds.
  • Proves that under finite-activity conditions (finitely many transitions per state), the ACOE holds globally.

Experimental results

Research questions

  • RQ1Can the existence of an optimal stationary policy in unbounded CTMDPs be established without constructing complex verifying functions?
  • RQ2To what extent can system stability results from queueing theory be leveraged to verify policy optimality in CTMDPs?
  • RQ3Under what conditions does the average-cost optimality equation (ACOE) hold in denumerable-state CTMDPs with unbounded rates and costs?
  • RQ4Is there a condition that simplifies the verification of optimal policies by reducing it to checking first-passage time and cost finiteness?

Key findings

  • The proposed condition for optimal stationary policy existence is based on the finiteness of expected first-passage time and expected cost to a reference state, which is easier to verify than prior function-based conditions.
  • The priority service (PS) policy in a two-queue system is proven to be a 0-standard policy, confirming its optimality under the new condition.
  • The expected cost from any state to the origin under the PS policy is finite, as shown via a Lyapunov function with exponential weighting and inequality constraints.
  • The average-cost optimality equation (ACOE) holds for all states when the policy is i₀-standard and only finitely many transitions are possible per state.
  • The method allows reuse of existing stability results from queueing theory to establish policy optimality without constructing problem-specific functions.
  • The result extends to cost functions that are increasing and polynomial in the state vector, broadening applicability.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.