Skip to main content
QUICK REVIEW

[论文解读] On the Convergence of Optimal Actions for Markov Decision Processes and the Optimality of $(s,S)$ Inventory Policies

Eugene A. Feinberg, Mark Edward Lewis|arXiv (Cornell University)|Jul 17, 2015
Supply Chain and Inventory Management参考文献 29被引用 3
一句话总结

本论文建立了在弱连续转移、无界成本及非紧动作集条件下,折扣成本与平均成本马尔可夫决策过程(MDPs)中最优动作的收敛性质。通过将这些结果应用于库存控制问题,证明了在一般需求分布下(无需假设离散或连续需求)(s,S) 库存策略的最优性,方法是展示当折扣因子趋近于1时最优阈值的收敛性。

ABSTRACT

This paper studies convergence properties of optimal values and actions for discounted and average-cost Markov Decision Processes (MDPs) with weakly continuous transition probabilities and applies these properties to the stochastic periodic-review inventory control problem with backorders, positive setup costs, and convex holding/backordering costs. The following results are established for MDPs with possibly noncompact action sets and unbounded cost functions: (i) convergence of value iterations to optimal values for discounted problems with possibly non-zero terminal costs, (ii) convergence of optimal finite-horizon actions to optimal infinite-horizon actions for total discounted costs, as the time horizon tends to infinity, and (iii) convergence of optimal discount-cost actions to optimal average-cost actions for infinite-horizon problems, as the discount factor tends to 1. Being applied to the setup-cost inventory control problem, the general results on MDPs imply the optimality of $(s,S)$ policies and convergence properties of optimal thresholds. In particular this paper analyzes the setup-cost inventory control problem without two assumptions often used in the literature: (a) the demand is either discrete or continuous or (b) the backordering cost is higher than the cost of backordered inventory if the amount of backordered inventory is large.

研究动机与目标

  • 在折扣成本与平均成本准则下,建立无限时域MDP中最优动作的一般收敛性质。
  • 通过将MDP的一般理论结果应用于随机库存问题,弥合MDP理论与库存控制之间的差距。
  • 在不假设需求为离散或连续分布的前提下,证明具有设置成本的库存系统中(s,S)策略的最优性。
  • 展示当时间范围增加或折扣因子趋近于1时,有限时域最优阈值(s_t, S_t)收敛到极限值(s, S)。
  • 通过去除常见假设(如有界持有成本或特定类型的需求分布)来扩展现有结果。

提出的方法

  • 利用弱连续转移概率和下确界紧致(inf-compact)成本函数,确保MDP中存在最优策略。
  • 应用带无界成本的折扣成本与平均成本MDP的值迭代及收敛定理。
  • 采用广义形式的最优方程与平均成本最优不等式(ACOI),推导结构化结果。
  • 利用价值函数的K-凸性,刻画在平均成本准则下的最优(s,S)策略。
  • 利用成本函数的连续性与凸性,证明有限时域最优动作收敛到无限时域极限。
  • 应用Feinberg与Liang(2018, 2019)关于价值函数连续性与K-凸性的结果于库存模型中。

实验结果

研究问题

  • RQ1在何种一般条件下,有限时域MDP中的最优动作会随着时域增加而收敛到无限时域最优动作?
  • RQ2当折扣因子趋近于1时,折扣MDP中的最优动作如何收敛到平均成本最优动作?
  • RQ3是否可以在不假设需求为离散或连续分布的前提下,证明(s,S)策略在库存控制问题中的最优性?
  • RQ4在有限时域设置成本库存问题中,什么条件能确保最优(s_t, S_t)阈值的收敛性?
  • RQ5在一般成本与转移假设下,平均成本最优方程是否以更强形式(如等式)成立?

主要发现

  • 值迭代收敛到折扣MDP的最优值,即使存在非零终端成本且成本无界。
  • 在总折扣成本准则下,当时间范围趋于无穷时,有限时域最优动作收敛到无限时域最优动作。
  • 当折扣因子趋近于1时,最优折扣成本动作收敛到最优平均成本动作。
  • 在一般需求分布下,(s,S)策略对无限时域设置成本库存控制问题是最优的,且无需假设需求为离散或连续分布。
  • 当时间范围增加时,最优阈值(s_t, S_t)收敛到极限值(s, S),其收敛性通过MDP收敛定理得以证明。
  • 在库存模型中,平均成本最优方程以更强形式(等式)成立,且价值函数具有K-凸性与连续性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。