Skip to main content
QUICK REVIEW

[論文レビュー] On Markov Decision Processes with Borel Spaces and an Average Cost Criterion

Huizhen Yu|arXiv (Cornell University)|Jan 10, 2019
Economic theories and models参考文献 37被引用数 4
ひとこと要約

本稿は、コンpact性や連続性を要件としない一般の条件下において、ボレル状態および行動空間を有するマルコフ決定過程(MDP)に対して、平均コスト最適性不等式(ACOI)を確立する。主に支配型条件とエゴロフの定理およびルジンの定理を用いることで、非有界かつ非負のコストモデルにおけるACOIの証明が可能となり、不連続なダイナミクスおよびコスト関数に対しても最適性の結果を得られる。

ABSTRACT

We consider average-cost Markov decision processes (MDPs) with Borel state and action spaces and universally measurable policies. For the nonnegative cost model and an unbounded cost model, we introduce a set of conditions under which we prove the average cost optimality inequality (ACOI) via the vanishing discount factor approach. Unlike most existing results on the ACOI, which require compactness/continuity conditions on the MDP, our result does not and can be applied to problems with discontinuous dynamics and one-stage costs. The key idea here is to replace the compactness/continuity conditions used in the prior work by what we call majorization type conditions. In particular, among others, we require that for each state, on selected subsets of actions at that state, the state transition stochastic kernel is majorized by finite measures, and we use this majorization property together with Egoroff's theorem to prove the ACOI. We also consider the minimum pair approach for average-cost MDPs and apply the majorization idea. For the case of a discrete action space and strictly unbounded costs, we prove the existence of a minimum pair that consists of a stationary policy and an invariant probability measure induced by the policy. This result is derived by combining Lusin's theorem with another majorization condition we introduce, and it can be applied to a class of countable action space MDPs in which, with respect to the state variable, the dynamics and one-stage costs are discontinuous.

研究の動機と目的

  • コンパクト性や連続性の条件を要件としない一般のボレル空間MDPに、平均コスト最適性理論を拡張すること。
  • 不連続なダイナミクスおよび1ステップコストを有するMDPにおける可測性および収束に関する問題を解決すること。
  • より弱い構造的仮定の下で、消える割引因子アプローチを用いてACOIを確立すること。
  • 非有界コストを有する離散的行動空間に対して、最小ペア(定常方策および不変測度)の存在を示すこと。
  • 従来の連続性およびコンパクト性要件に代わる支配型条件に基づく枠組みを構築すること。

提案手法

  • 特定の行動部分集合上で、状態遷移核が有限測度によって支配される支配型条件を導入する。
  • エゴロフの定理を用いて、関数が一様収束性を示す大測度の集合を抽出する。
  • ルジンの定理を用いて、最小ペアの存在証明における方策の可測選択を構築する。
  • 消える割引因子アプローチを用い、割引コスト最適性方程式の極限を取ることでACOIを導出する。
  • 支配型条件とライアプノフ型条件を組み合わせ、非有界コストモデルにおける可積分性および安定性を保証する。
  • 可測選択定理(例:[2, Prop. 7.50])を用いて、部分最適行動から普遍可測方策を構築する。

実験結果

リサーチクエスチョン

  • RQ1コンパクト性や連続性の仮定を要件としないボレル空間MDPにおいて、平均コスト最適性不等式(ACOI)を確立できるか?
  • RQ2連続性およびコンパクト性を代替するためのどのような条件が、平均コストMDPにおける収束性および最適性を保証できるか?
  • RQ3非有界コストを有するMDPにおいて、最小ペア(定常方策および不変測度)が存在する条件は何か?
  • RQ4エゴロフの定理およびルジンの定理をどのように活用することで、ダイナミクスおよびコスト関数の不連続性に対処できるか?
  • RQ5支配型条件の下で、消える割引因子アプローチをACOIの証明に適応できるか?

主な発見

  • 非負コストモデルおよびライアプノフ型条件を満たす非有界コストモデルに対して、連続性やコンパクト性を要件としないACOIが確立された。
  • 証明は、行動部分集合上で遷移核が有限測度によって支配される支配型条件に依拠しており、エゴロフの定理の適用を可能にする。
  • 厳密に非有界なコストを有する離散的行動空間に対して、定常方策および不変測度からなる最小ペアが存在することが、ルジンの定理および新たな支配型条件を用いて証明された。
  • 構築された非確率的マルコフ方策は平均コスト最適であり、ACOIのもとで定常方策がε-最適であることが示された。
  • 導入された支配型条件およびライアプノフ条件のもとで、価値関数およびコスト到達関数が有限かつ可測であることが示された。
  • 重み付きノルムにおける割引コスト作用素の収縮性が確立され、価値反復プロセスの収束が保証された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。