[论文解读] Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
本文提出了一种针对未知马尔可夫跳变线性系统(MJS)的自适应控制框架,结合系统辨识与确定等价控制。该研究建立了系统辨识的最优 $Ø(1/ar{\sqrt{T}})$ 样本复杂度,并在均方稳定条件下实现了 $Ø(\sqrt{T})$ 的遗憾边界,若具备部分系统知识,则可进一步优化至 $\mathrm{polylog}(T)$,其新颖的分析方法专为处理 MJS 的混合动态特性与较弱的稳定性概念而设计。
Learning how to effectively control unknown dynamical systems is crucial for intelligent autonomous systems. This task becomes a significant challenge when the underlying dynamics are changing with time. Motivated by this challenge, this paper considers the problem of controlling an unknown Markov jump linear system (MJS) to optimize a quadratic objective. By taking a model-based perspective, we consider identification-based adaptive control of MJSs. We first provide a system identification algorithm for MJS to learn the dynamics in each mode as well as the Markov transition matrix, underlying the evolution of the mode switches, from a single trajectory of the system states, inputs, and modes. Through martingale-based arguments, sample complexity of this algorithm is shown to be $\mathcal{O}(1/\sqrt{T})$. We then propose an adaptive control scheme that performs system identification together with certainty equivalent control to adapt the controllers in an episodic fashion. Combining our sample complexity results with recent perturbation results for certainty equivalent control, we prove that when the episode lengths are appropriately chosen, the proposed adaptive control scheme achieves $\mathcal{O}(\sqrt{T})$ regret, which can be improved to $\mathcal{O}(polylog(T))$ with partial knowledge of the system. Our proof strategy introduces innovations to handle Markovian jumps and a weaker notion of stability common in MJSs. Our analysis provides insights into system theoretic quantities that affect learning accuracy and control performance. Numerical simulations are presented to further reinforce these insights.
研究动机与目标
- 为解决未知马尔可夫跳变线性系统(MJS)在切换动态下缺乏非渐近学习保证的问题。
- 设计一种系统辨识算法,可在均方稳定条件下,仅从单一轨迹中估计模式相关的动态特性与马尔可夫转移矩阵。
- 设计一种自适应控制方案,将系统辨识与确定等价控制相结合,用于分段学习。
- 为自适应控制策略建立紧致的遗憾边界,证明其在轨迹长度 $T$ 上的最优性。
- 提供关于影响 MJS 中学习与控制性能的系统理论量(如稳定性裕度、混合时间等)的理论洞见。
提出的方法
- 提出一种系统辨识算法(算法 1),可从状态、输入与模式的单一轨迹中估计 $s$ 个模式相关的状态-输入矩阵 $(\mathbf{A}_i, \mathbf{B}_i)$ 与马尔可夫转移矩阵 $\mathbf{T}$。
- 利用混合时间论证,推导出样本复杂度为 $\mathcal{O}((n+p)\log T \sqrt{s/T})$,该复杂度在对数因子内达到最优。
- 引入一种分段自适应控制方案,每段内同时执行系统辨识与控制更新,采用确定等价控制策略。
- 采用时变分段长度策略,其中 $T_i = \gamma T_{i-1}$,以平衡探索与利用。
- 结合高概率浓度不等式与扰动理论,推导出在均方稳定条件下的遗憾边界。
- 提出一种新颖的证明策略,以处理 MJS 的混合动态特性与较弱的均方稳定概念(该概念不保证几乎必然收敛)。
实验结果
研究问题
- RQ1从单一轨迹中识别未知马尔可夫跳变线性系统动态特性的最优样本复杂度是什么?
- RQ2在均方稳定条件下,自适应控制策略是否能在马尔可夫切换动态下实现次线性遗憾?
- RQ3系统理论量(如混合时间、谱半径与稳定性裕度)如何影响 MJS 中的学习与控制性能?
- RQ4若具备部分系统知识,能否将遗憾边界从 $\mathcal{O}(\sqrt{T})$ 改进至 $\mathcal{O}(\mathrm{polylog}(T))$?
- RQ5由于缺乏确定性稳定性与模式切换的存在,分析 MJS 的学习与控制面临哪些关键挑战?
主要发现
- 所提系统辨识算法实现了 $\mathcal{O}((n+p)\log T \sqrt{s/T})$ 的样本复杂度,对轨迹长度 $T$ 的依赖为 $\mathcal{O}(1/\sqrt{T})$,在对数因子内达到最优。
- 所提出的自适应控制方案在均方稳定条件下实现了 $\mathcal{O}(\sqrt{T})$ 的遗憾边界,且以高概率 $1 - \delta$ 成立。
- 在具备部分系统知识的条件下,遗憾边界可提升至 $\mathcal{O}(\mathrm{polylog}(T))$,表明先验信息具有显著优势。
- 分析表明,稳定性裕度 $\bar{\theta}$、噪声方差 $\sigma_{\mathbf{w}}^2$ 与最小平稳概率 $\pi_{\min}$ 显著影响遗憾的缩放特性。
- 该证明框架成功处理了 MJS 的混合动态特性与非确定性稳定性,首次为该类系统提供了非渐近遗憾保证。
- 数值仿真验证了理论洞见,表明自适应控制器可收敛至最优性能,且遗憾边界与理论预测一致。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。