Skip to main content
QUICK REVIEW

[論文レビュー] On the Convergence of Optimal Actions for Markov Decision Processes and the Optimality of $(s,S)$ Inventory Policies

Eugene A. Feinberg, Mark Edward Lewis|arXiv (Cornell University)|Jul 17, 2015
Supply Chain and Inventory Management参考文献 29被引用数 3
ひとこと要約

この論文は、弱く連続な遷移、非有界コスト、非コンパクトな行動集合を伴う割引コストおよび平均コストのマルコフ決定過程(MDP)における最適行動の収束特性を確立する。これらの結果を応用して、離散的または連続的であるという仮定を必要としない一般の需要分布のもとで、(s,S)在庫政策の最適性を示す。これは、割引係数が1に近づく際の最適しきい値の収束を示すことによって達成される。

ABSTRACT

This paper studies convergence properties of optimal values and actions for discounted and average-cost Markov Decision Processes (MDPs) with weakly continuous transition probabilities and applies these properties to the stochastic periodic-review inventory control problem with backorders, positive setup costs, and convex holding/backordering costs. The following results are established for MDPs with possibly noncompact action sets and unbounded cost functions: (i) convergence of value iterations to optimal values for discounted problems with possibly non-zero terminal costs, (ii) convergence of optimal finite-horizon actions to optimal infinite-horizon actions for total discounted costs, as the time horizon tends to infinity, and (iii) convergence of optimal discount-cost actions to optimal average-cost actions for infinite-horizon problems, as the discount factor tends to 1. Being applied to the setup-cost inventory control problem, the general results on MDPs imply the optimality of $(s,S)$ policies and convergence properties of optimal thresholds. In particular this paper analyzes the setup-cost inventory control problem without two assumptions often used in the literature: (a) the demand is either discrete or continuous or (b) the backordering cost is higher than the cost of backordered inventory if the amount of backordered inventory is large.

研究の動機と目的

  • 割引コストおよび平均コスト基準の下で、無限時間ホライズンMDPにおける最適行動の一般化された収束特性を確立すること。
  • MDP理論と在庫制御の間のギャップを埋めるために、一般MDPの結果を確率的在庫問題に応用すること。
  • 離散的または連続的需要の仮定をせず、一般の需要分布のもとで、セットアップコストを伴う在庫システムにおける(s,S)政策の最適性を証明すること。
  • 時間ホライズンが延長されるか、割引係数が1に近づくにつれて、最適しきい値(s_t, S_t)が限界値(s, S)に収束することを示すこと。
  • 有界な保持コストや特定の需要分布型という一般的な仮定を排除することで、既存の結果を拡張すること。

提案手法

  • 弱く連続な遷移確率と下限コンパクトなコスト関数を用いて、MDPにおける最適方策の存在を保証する。
  • 非有界コストを伴う割引コストおよび平均コストMDPにおける価値反復と収束定理を適用する。
  • 一般化された形での最適性方程式と平均コスト最適性不等式(ACOI)を用いて構造的結果を導出する。
  • 平均コスト基準における最適(s,S)方策を特徴付けるために、価値関数のK-凸性を活用する。
  • コスト関数の連続性と凸性を用いて、有限時系列の最適行動が無限時系列の極限に収束することを証明する。
  • FeinbergとLiang(2018, 2019)の価値関数の連続性およびK-凸性に関する結果を、在庫モデルに応用する。

実験結果

リサーチクエスチョン

  • RQ1有限時系列MDPにおける最適行動は、時間ホライズンが延長される際に、無限時系列の最適行動にどのような一般条件で収束するか?
  • RQ2割引係数が1に近づくにつれて、割引コスト最適行動は平均コスト最適行動にどのように収束するか?
  • RQ3(s,S)方策は、離散的または連続的需要分布の仮定をしない在庫制御問題において最適であると証明できるか?
  • RQ4有限時系列のセットアップコスト在庫問題における最適(s_t, S_t)しきい値の収束を保証する条件は何か?
  • RQ5一般のコストおよび遷移仮定のもとで、平均コスト最適性方程式はより強い形(等式)で成立するか?

主な発見

  • 価値反復は、可能に非ゼロの終端コストと非有界コストを伴う割引MDPにおいて、最適値に収束する。
  • 時間ホライズンが無限大に近づくにつれて、有限時系列の最適行動は、総割引コスト基準における無限時系列の最適行動に収束する。
  • 割引係数が1に近づくにつれて、最適割引コスト行動は、最適平均コスト行動に収束する。
  • 一般の需要分布のもとで、無限時系列のセットアップコスト在庫制御問題において(s,S)方策が最適である。離散的または連続的需要の仮定は不要である。
  • 時間ホライズンが延長されるにつれて、最適しきい値(s_t, S_t)は限界値(s, S)に収束し、MDP収束定理を用いてその収束を証明した。
  • 平均コスト最適性方程式は、より強い形(等式)で成立し、価値関数はK-凸かつ連続である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。