Skip to main content
QUICK REVIEW

[论文解读] Regret-Optimal Full-Information Control.

Oron Sabag, Gautam Goel|arXiv (Cornell University)|May 4, 2021
Advanced Bandit Algorithms Research参考文献 21被引用 10
一句话总结

本文提出了一种后悔最优的全信息控制器,通过将问题简化为Nehari逼近,最小化最坏情况下的后悔值——即因果控制器与具有未来信息的非因果(先知)控制器之间LQR代价的差异。最优控制器被证明是标准H₂状态反馈律与通过求解两个李雅普诺夫方程得到的有限维控制器之和,从而在H₂与H∞性能之间实现平衡,并对未来的扰动具有鲁棒性。

ABSTRACT

We consider the infinite-horizon, discrete-time full-information control problem. Motivated by learning theory, as a criterion for controller design we focus on regret, defined as the difference between the LQR cost of a causal controller (that has only access to past and current disturbances) and the LQR cost of a clairvoyant one (that has also access to future disturbances). In the full-information setting, there is a unique optimal non-causal controller that in terms of LQR cost dominates all other controllers. Since the regret itself is a function of the disturbances, we consider the worst-case regret over all possible bounded energy disturbances, and propose to find a causal controller that minimizes this worst-case regret. The resulting controller has the interpretation of guaranteeing the smallest possible regret compared to the best non-causal controller, no matter what the future disturbances are. We show that the regret-optimal control problem can be reduced to a Nehari problem, i.e., to approximate an anticausal operator with a causal one in the operator norm. In the state-space setting, explicit formulas for the optimal regret and for the regret-optimal controller (in both the causal and the strictly causal settings) are derived. The regret-optimal controller is the sum of the classical $H_2$ state-feedback law and a finite-dimensional controller obtained from the Nehari problem. The controller construction simply requires the solution to the standard LQR Riccati equation, in addition to two Lyapunov equations. Simulations over a range of plants demonstrates that the regret-optimal controller interpolates nicely between the $H_2$ and the $H_\infty$ optimal controllers, and generally has $H_2$ and $H_\infty$ costs that are simultaneously close to their optimal values. The regret-optimal controller thus presents itself as a viable option for control system design.

研究动机与目标

  • 为解决标准H₂与H∞控制器的局限性,提出一种基于全信息控制中后悔的新性能准则。
  • 在所有有界能量扰动下最小化最坏情况下的后悔值,确保对任何未来扰动实现方式均具有鲁棒性。
  • 推导因果与严格因果设置下后悔最优控制器的显式状态空间公式。
  • 证明后悔最优控制器在H₂与H∞性能之间实现插值,两种度量下的代价均接近最优。

提出的方法

  • 将后悔定义为因果控制器的LQR代价与具有未来扰动信息的非因果(先知)控制器代价之间的差值。
  • 将后悔最优控制问题转化为Nehari问题:在算子范数下,用因果算子逼近一个反因果算子。
  • 推导出后悔最优控制器的显式状态空间表达式,即标准H₂状态反馈增益与来自Nehari解的有限维控制器之和。
  • 求解标准LQR黎卡提方程以及两个附加的李雅普诺夫方程,以构建后悔最优控制器。
  • 利用Nehari问题的解,确保控制器在所有有界能量扰动序列下最小化最坏情况下的后悔值。
  • 通过在一系列被控对象上的仿真验证控制器性能,并比较H₂与H∞代价。

实验结果

研究问题

  • RQ1能否设计出一种因果控制器,使其在全信息设置下相对于具有未来信息的非因果控制器,最小化最坏情况下的后悔值?
  • RQ2与经典的H₂与H∞控制器相比,后悔最优控制器在性能权衡方面表现如何?
  • RQ3后悔最优控制器的显式状态空间结构是什么?如何实现高效计算?
  • RQ4后悔最优控制问题能否被简化为如Nehari问题这样的经典算子逼近问题?
  • RQ5后悔最优控制器是否能在H₂与H∞代价上同时实现接近最优的平衡性能?

主要发现

  • 后悔最优控制器被显式构造为标准H₂状态反馈律与通过求解两个李雅普诺夫方程得到的有限维控制器之和。
  • 最优后悔值由Nehari问题的解决定,确保在所有有界能量扰动下最坏情况下的后悔值最小化。
  • 控制器的构造仅需求解标准LQR黎卡提方程以及两个附加的李雅普诺夫方程,从而实现高效计算。
  • 仿真结果表明,后悔最优控制器在H₂与H∞代价上均接近各自最优值,表明其具有有利的性能权衡。
  • 后悔最优控制器在H₂与H∞最优控制器之间实现平滑插值,为实际控制系统设计提供了一种鲁棒的替代方案。
  • 该方法在后悔最小化与经典Nehari问题之间建立了正式联系,为鲁棒控制设计提供了新的理论框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。