[論文レビュー] Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
本稿は、システム同定とCertainty Equivalent制御を用いて、未知のマーケフ跳躍線形システム(MJS)のための適応制御フレームワークを提案する。システム同定に対しては、最適な $Ø(1/ar{\sqrt{T}})$ のサンプル複雑度を確立し、平均二乗安定性のもとで $Ø(\sqrt{T})$ のレグレットバウンドを達成する。部分的なシステム知識があると、$\mathrm{polylog}(T)$ に改善され、MJSのハイブリッドダイナミクスと弱い安定性概念に特化した新しい解析が用いられる。
Learning how to effectively control unknown dynamical systems is crucial for intelligent autonomous systems. This task becomes a significant challenge when the underlying dynamics are changing with time. Motivated by this challenge, this paper considers the problem of controlling an unknown Markov jump linear system (MJS) to optimize a quadratic objective. By taking a model-based perspective, we consider identification-based adaptive control of MJSs. We first provide a system identification algorithm for MJS to learn the dynamics in each mode as well as the Markov transition matrix, underlying the evolution of the mode switches, from a single trajectory of the system states, inputs, and modes. Through martingale-based arguments, sample complexity of this algorithm is shown to be $\mathcal{O}(1/\sqrt{T})$. We then propose an adaptive control scheme that performs system identification together with certainty equivalent control to adapt the controllers in an episodic fashion. Combining our sample complexity results with recent perturbation results for certainty equivalent control, we prove that when the episode lengths are appropriately chosen, the proposed adaptive control scheme achieves $\mathcal{O}(\sqrt{T})$ regret, which can be improved to $\mathcal{O}(polylog(T))$ with partial knowledge of the system. Our proof strategy introduces innovations to handle Markovian jumps and a weaker notion of stability common in MJSs. Our analysis provides insights into system theoretic quantities that affect learning accuracy and control performance. Numerical simulations are presented to further reinforce these insights.
研究の動機と目的
- スイッチングダイナミクスを有する未知のマーケフ跳躍線形システム(MJS)に対して、非漸近的学習保証が不足している問題に対処すること。
- 平均二乗安定性のもとで、1つの軌道からのモード固有のダイナミクスとマーケフ遷移行列を推定するシステム同定アルゴリズムを開発すること。
- エピソード的学習を実現するため、システム同定とCertainty Equivalent制御を統合した適応制御方式を設計すること。
- 適応制御方策のタイトなレグレットバウンドを確立し、軌道長 $T$ における最適性を示すこと。
- MJSにおける学習と制御性能に影響を与えるシステム理論的量(安定性マージン、混合時間、スペクトル半径など)に関する理論的洞察を提供すること。
提案手法
- 1つの軌道からの状態、入力、モードの観測値を用いて、$s$ 個のモード固有の状態入力行列 $(\mathbf{A}_i, \mathbf{B}_i)$ とマーケフ遷移行列 $\mathbf{T}$ を推定するシステム同定アルゴリズム(アルゴリズム1)を提案する。
- 混合時間の議論を用いて、$\mathcal{O}((n+p)\log T \sqrt{s/T})$ のサンプル複雑度を導出する。これは対数要因を除いて最適である。
- エピソード的適応制御方式を導入し、各エピソードでシステム同定と制御更新を実行する。制御にはCertainty Equivalent制御を用いる。
- 探索と活用のバランスを取るために、$T_i = \gamma T_{i-1}$ の時間依存的エピソード長ポリシーを採用する。
- 高確率の濃度不等式と摂動理論を用いて、平均二乗安定性のもとでレグレットバウンドを導出する。
- MJSのハイブリッド性と、確実収束を保証しない弱い平均二乗安定性という性質に対処するための、新しい証明戦略を導入する。
実験結果
リサーチクエスチョン
- RQ1未知のマーケフ跳躍線形システムのダイナミクスを、1つの軌道から同定するための最適なサンプル複雑度は何か?
- RQ2平均二乗安定性のもとで、マーケフ的スイッチングダイナミクスが存在する状況において、適応制御方策がサブラインアーなレグレットを達成できるか?
- RQ3混合時間、スペクトル半径、安定性マージンといったシステム理論的量は、MJSにおける学習と制御性能にどのように影響を与えるか?
- RQ4部分的なシステム知識があると、レグレットバウンドを $\mathcal{O}(\sqrt{T})$ から $\mathcal{O}(\mathrm{polylog}(T))$ に改善できるか?
- RQ5決定的安定性の欠如とモードスイッチングの存在により、MJSの学習と制御の解析において直面する主な課題は何か?
主な発見
- システム同定アルゴリズムは、$\mathcal{O}((n+p)\log T \sqrt{s/T})$ のサンプル複雑度を達成し、軌道長 $T$ に対して $\mathcal{O}(1/\sqrt{T})$ の依存性を示す。これは対数要因を除いて最適である。
- 提案された適応制御方式は、平均二乗安定性のもとで、高確率 $1 - \delta$ で $\mathcal{O}(\sqrt{T})$ のレグレットバウンドを達成する。
- 部分的なシステム知識があると、レグレットバウンドは $\mathcal{O}(\mathrm{polylog}(T))$ に改善され、事前情報の利点が明確に示される。
- 解析により、安定性マージン $\bar{\theta}$、ノイズ分散 $\sigma_{\mathbf{w}}^2$、最小定常確率 $\pi_{\min}$ がレグレットスケーリングに顕著に影響することが判明した。
- 証明フレームワークは、MJSのハイブリッドダイナミクスと非決定的安定性を効果的に扱い、このクラスのシステムに対する非漸近的レグレット保証を初めて得た。
- 数値シミュレーションにより理論的洞察が検証され、適応制御器が最適性能に収束し、予測通りにレグレットがスケーリングすることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。