[论文解读] On Optimality of Adaptive Linear-Quadratic Regulators.
本文通过引入自适应策略的新型分解,对自适应线性二次调节器中的遗憾提供了精确表征,证明了修正的确定性等价方案可实现近乎平方根阶的遗憾率。它建立了最优参数辨识速率,并确定了将遗憾降至对数阶所需的最少附加信息,从而解决了自适应性理论中的关键空白。
Adaptive regulation of linear systems represents a canonical problem in stochastic control. Performance of adaptive control policies is assessed through the regret with respect to the optimal regulator, that reflects the increase in the operating cost due to uncertainty about the parameters that drive the dynamics of the system. However, available results in the literature do not provide a sharp quantitative characterization of the effect of the unknown dynamics parameters on the regret. Further, there are issues on how easy it is to implement the adaptive policies proposed in the literature. Finally, results regarding the accuracy that the system's parameters are identified are scarce and rather incomplete. This study aims to comprehensively address these three issues. First, by introducing a novel decomposition of adaptive policies, we establish a sharp expression for the regret of an arbitrary policy in terms of the deviations from the optimal regulator. Second, we show that adaptive policies based on a slight modification of the widely used Certainty Equivalence scheme are optimal. Specifically, we establish a regret of (nearly) square-root rate for two families of randomized adaptive policies. The presented regret bounds are obtained by using anti-concentration results on the random matrices employed when randomizing the estimates of the unknown dynamics parameters. Moreover, we study the minimal additional information needed on dynamics matrices for which the regret will become of logarithmic order. Finally, the rate at which the unknown parameters of the system are being identified is specified for the proposed adaptive policies.
研究动机与目标
- 为由于系统参数未知而引起的自适应线性二次调节器中的遗憾提供精确且量化的表征。
- 解决现有自适应控制策略在随机控制设置下的实用性和可实现性问题。
- 确定在所提出的自适应策略下,未知系统参数的辨识速率。
- 确定将遗憾降低至对数阶所需的最少附加信息。
- 填补文献中关于遗憾边界、策略可实现性以及参数估计精度方面的空白。
提出的方法
- 提出一种自适应策略的新型分解,将遗憾表示为与最优调节器偏离程度的函数。
- 分析带有随机参数估计的修正确定性等价策略,以实现最优遗憾性能。
- 应用随机矩阵的反集中度结果,推导出随机自适应策略的紧致遗憾边界。
- 使用矩阵集中不等式,量化随机化下参数估计的统计行为。
- 推导出当额外先验信息满足何种条件时,遗憾将从平方根阶过渡到对数阶。
- 利用统计学习理论,制定所提出自适应策略的参数辨识速率。
实验结果
研究问题
- RQ1在自适应LQR中,策略与最优性之间的偏离程度与遗憾之间存在何种精确关系?
- RQ2修正的确定性等价策略是否可在自适应线性二次控制中实现近乎最优的遗憾率?
- RQ3在所提出的自适应策略下,未知系统参数的估计速率是多少?
- RQ4何种最少附加信息可使遗憾降低至对数阶?
- RQ5随机化以及随机矩阵的反集中度性质如何影响遗憾边界?
主要发现
- 任何自适应策略的遗憾均通过一种分解被精确表征,该分解量化了其与最优调节器的偏离程度。
- 修正的确定性等价策略实现了近乎时间跨度平方根阶的遗憾率,与基本下界一致。
- 随机矩阵的反集中度性质在推导随机估计方案的紧致遗憾边界中起到了关键作用。
- 所提出的自适应策略以与线性回归模型中统计最优性一致的速率识别系统参数。
- 当获得关于动态矩阵的额外结构信息时,遗憾可被降低至对数阶。
- 本研究首次在统一框架下完整表征了自适应LQR的遗憾、参数估计速率及可实现性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。