Skip to main content
QUICK REVIEW

[论文解读] On the Convergence of Optimal Actions for Markov Decision Processes and the Optimality of $(s,S)$ Policies for Inventory Control

Eugene A. Feinberg, Mark Edward Lewis|arXiv (Cornell University)|Jul 17, 2015
Supply Chain and Inventory Management被引用 8
一句话总结

本文建立了在弱连续转移、无界成本和非紧致动作集条件下,折扣成本与平均成本马尔可夫决策过程(MDPs)中最优动作的收敛性。证明了有限-horizon最优动作向无限-horizon动作的收敛性,以及折扣成本动作向平均成本动作的收敛性,并将这些结果应用于库存控制问题,证明了(s,S)策略的最优性,而无需假设需求为离散或连续,也无需假设高缺货成本惩罚。

ABSTRACT

This paper describes results on the existence of optimal policies and convergence properties of optimal actions for discounted and average-cost Markov Decision Processes (MDPs) with weakly continuous transition probabilities. It is possible that cost functions are unbounded and action sets are not compact. The following results are established for such MDPs: (i) convergence of value iterations to optimal values for discounted problems with possibly non-zero terminal costs, (ii) convergence of optimal finite-horizon actions to optimal infinite-horizon actions for total discounted costs, as the time horizon tends to infinity, and (iii) convergence of optimal discount-cost actions to optimal average-cost actions for infinite-horizon problems, as the discount factor tends to 1. The general results on MDPs are applied to the classic stochastic periodic-review inventory control problems with backorders, for which they imply the optimality of $(s,S)$ policies and convergence properties of optimal thresholds. In particular we analyze inventory control problems without two assumptions often used in the literature: (a) the demand is either discrete or continuous or (b) backordering is more expensive that the cost of backordered inventory if the backordered amount is large.

研究动机与目标

  • 建立具有弱连续转移概率的折扣成本与平均成本MDPs中最优动作收敛性的理论基础。
  • 分析具有无界成本函数和非紧致动作集的MDPs,放宽文献中常见的假设。
  • 将一般MDP收敛结果应用于随机周期审查库存控制问题。
  • 在无需假设需求为离散或连续分布的前提下,证明(s,S)策略在库存控制中的最优性。
  • 消除对大缺货量下缺货成本超过持有成本的假设。

提出的方法

  • 使用值迭代方法,证明在可能具有非零终端成本的折扣MDPs中,值迭代收敛到最优值。
  • 应用转移概率的弱连续性,确保随着时域增加,最优有限-horizon动作收敛到无限自horizon动作。
  • 建立最优折扣成本动作在折扣因子趋近于1时收敛到最优平均成本动作。
  • 采用动态规划技术和连续性论证,处理无界成本函数和非紧致动作集。
  • 将一般MDP收敛结果应用于具有缺货的周期审查库存模型的特定结构。
  • 证明在推导出的条件下,(s,S)策略为最优,即使在缺乏标准需求或成本假设的情况下亦成立。

实验结果

研究问题

  • RQ1在时域趋于无穷时,MDPs中有限-horizon最优动作在何种条件下收敛到无限自horizon最优动作?
  • RQ2当折扣因子趋近于1时,折扣MDPs中的最优动作如何收敛到平均成本最优动作?
  • RQ3是否可以在不假设需求分布为离散或连续的前提下,证明(s,S)策略在库存控制模型中的最优性?
  • RQ4当缺货成本不必然超过大缺货量下的持有成本时,(s,S)策略的最优性是否仍然成立?
  • RQ5在弱连续性和无界成本条件下,何种MDP一般条件可确保值迭代和最优动作的收敛性?

主要发现

  • 对于具有无界成本和非零终端成本的折扣MDPs,值迭代收敛到最优值。
  • 随着时域趋于无穷,最优有限自horizon动作收敛到总折扣成本下的最优无限自horizon动作。
  • 当折扣因子趋近于1时,最优折扣成本动作收敛到最优平均成本动作。
  • 在具有缺货的周期审查库存控制中,(s,S)策略为最优,即使不假设需求为离散或连续。
  • (s,S)策略的最优性在无需缺货成本超过持有成本的假设下依然成立,即使在大缺货量下亦然。
  • 结果可推广至具有弱连续转移概率、无界成本函数和非紧致动作集的MDPs。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。