Skip to main content
QUICK REVIEW

[論文レビュー] Online Markov Decision Processes with Terminal Law Constraints

Bianca Marin Moreno, Margaux Brégère|arXiv (Cornell University)|Jan 12, 2026
Advanced Bandit Algorithms Research被引用数 0
ひとこと要約

この論文は、未知のダイナミクスと敵対的損失を伴うオンラインMDPに対するリセットなしの周期フレームワークを提案し、周期方針を定義し、マルチエージェント設定でのサブ線形の周期レグレットを実現するアルゴリズムを提供する。

ABSTRACT

Traditional reinforcement learning usually assumes either episodic interactions with resets or continuous operation to minimize average or cumulative loss. While episodic settings have many theoretical results, resets are often unrealistic in practice. The infinite-horizon setting avoids this issue but lacks non-asymptotic guarantees in online scenarios with unknown dynamics. In this work, we move towards closing this gap by introducing a reset-free framework called the periodic framework, where the goal is to find periodic policies: policies that not only minimize cumulative loss but also return the agents to their initial state distribution after a fixed number of steps. We formalize the problem of finding optimal periodic policies and identify sufficient conditions under which it is well-defined for tabular Markov decision processes. To evaluate algorithms in this framework, we introduce the periodic regret, a measure that balances cumulative loss with the terminal law constraint. We then propose the first algorithms for computing periodic policies in two multi-agent settings and show they achieve sublinear periodic regret of order $ ilde O(T^{3/4})$. This provides the first non-asymptotic guarantees for reset-free learning in the setting of $M$ homogeneous agents, for $M > 1$.

研究の動機と目的

  • 未知のダイナミクスを持つリセットフリーオンラインMDPにおける周期方針を形式化する。
  • 終端分布制約を考慮した周期レグレット指標を提案する。
  • マルチエージェント・敵対的設定における周期方針を計算するアルゴリズムを開発する。
  • M>1エージェントシナリオにおける周期レグレットの非漸近的保証を確立する。

提案手法

  • rho P_pi = rho を満たす周期方針を定義し、Nステップ後に初期分布へ戻ることを保証する。
  • 状態-行動分布上の一般 convex 損失を扱うための convex RL フレームワークを導入する。
  • 未知転移と敵対的損失に対処するボーナスベースの探索ミラー降下法(MDPP-K)を開発する。
  • ボーナスを用いてフェアな不等式へ変換された終端法則制約を含む制約付きMDPの定式化。
  • 2つのフレームワークを提供:フレームワーク1は既知の rho_t、フレームワーク2は推定 rho_t と制限リセットで、いずれも周期レグレット境界をもたらす。
  • 適切な条件の下で tilde-O(T^{3/4}) の非漸近的周期レグレット境界を分析・証明する。

実験結果

リサーチクエスチョン

  • RQ1未知のダイナミクスを持つオンラインMDPにおける周期方針の存在条件は何か。
  • RQ2累積損失と終端分布制約のバランスを取る周期レグレットをどう定義・最小化するか。
  • RQ3複数の均質エージェントを前提としたリセットなし学習に対して、証明可能な非漸近的保証を備えたオンラインアルゴリズムを設計できるか。
  • RQ4未知転移ダイナミクスは実現性にどのような影響を与え、ボーナスは探索と制約充足をどう保証するか。

主な発見

  • 周期方針の概念と rho へ収束させる収縮仮定 2 を提案し、遍在性を確保。
  • 周期レグレット R_T を提案し、累積損失の差と終端分布の偏差を組み合わせる。
  • MDPP-Kアルゴリズムは M>1 のエージェントに対して tilde-O(T^{3/4}) のサブ線形周期レグレットを達成。
  • 既知 rho_t と未知 rho_t(リセット制限付き)という2つのフレームワークと、それぞれのレグレット保証を提供。
  • 敵対的損失下の制約付きMDP問題に対する実現性と高確率境界を示す。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。