Skip to main content
QUICK REVIEW

[论文解读] Measurement-Feedback Control with Optimal Data-Dependent Regret

Gautam Goel, Babak Hassibi|arXiv (Cornell University)|Sep 14, 2022
Advanced Bandit Algorithms Research被引用 4
一句话总结

本文提出了一类针对线性系统的后悔最优测量反馈控制器,通过将问题转化为合成系统中的H∞最优控制,实现了最优数据依赖性后悔。推导出两种控制器:一种对驱动干扰能量和测量干扰能量的联合依赖达到最优,另一种对驱动干扰路径长度和测量干扰能量的联合依赖达到最优,并证明了一般情况下有界竞争比是不可能实现的。

ABSTRACT

Inspired by online learning, data-dependent regret has recently been proposed as a criterion for controller design. In the regret-optimal control paradigm, causal controllers are designed to minimize regret against a hypothetical optimal noncausal controller, which selects the globally cost-minimizing sequence of control actions given noncausal access to the disturbance sequence. We extend regret-optimal control to the more challenging measurement-feedback setting, where the online controller must compete against the optimal noncausal controller without directly observing the state or the driving disturbance. We show that no measurement-feedback controller can have bounded competitive ratio or regret which is bounded by the pathlength of the measurement disturbance. We do derive, however, a controller whose regret has optimal dependence on the joint energy of the driving and measurement disturbances, and another controller whose regret has optimal dependence on the pathlength of the driving disturbance and the energy of the measurement disturbance. The key technique we introduce is a reduction from regret-optimal measurement-feedback control to $H_{\infty}$-optimal measurement-feedback control in a synthetic system. We present numerical simulations which illustrate the efficacy of our proposed control algorithms.

研究动机与目标

  • 解决在测量反馈设置下设计因果控制器以最小化与非因果最优控制器相比的后悔问题。
  • 克服在一般线性系统中,测量反馈下有界竞争比或路径长度有界后悔不可行的局限性。
  • 开发在干扰复杂度度量(联合能量与混合路径长度-能量)上具有最优后悔依赖的控制器。
  • 建立从后悔最优测量反馈控制到合成系统中H∞最优控制的约化框架,以实现可计算的设计。

提出的方法

  • 将后悔最优测量反馈控制问题表述为相对于全知非因果控制器的后悔最小化问题。
  • 引入一个合成系统,将原始的测量反馈问题转化为H∞最优控制问题。
  • 通过状态空间变换,将原始系统动态和干扰模型嵌入到具有增广状态和控制输入的合成系统中。
  • 在合成系统上应用H∞最优控制框架,推导出具有可证明最优后悔边界的控制器。
  • 推导出两种不同的控制器:一种最小化与驱动和测量干扰联合能量最优依赖的后悔,另一种实现与驱动干扰路径长度和测量干扰能量最优权衡的后悔。
  • 利用分离原理和Riccati方程解,确保所推导控制器的稳定性和最优性。

实验结果

研究问题

  • RQ1测量反馈控制器能否实现有界竞争比或与测量干扰路径长度相关的有界后悔?
  • RQ2在不同干扰复杂度度量下,测量反馈控制器可实现的最优后悔依赖是什么?
  • RQ3能否将后悔最优测量反馈控制约化为合成系统中的H∞最优控制问题?
  • RQ4所提出的控制器在不同干扰类型下与标准H₂和H∞控制器相比性能如何?
  • RQ5后悔最小化在测量反馈控制设置下的基本限制是什么?

主要发现

  • 不存在测量反馈控制器能够实现有界竞争比或与测量干扰路径长度相关的有界后悔。
  • 所提出的控制器实现了与驱动和测量干扰联合能量最优依赖的后悔,达到理论下界。
  • 第二种控制器实现了与驱动干扰路径长度和测量干扰能量最优依赖的后悔。
  • 在双积分器系统上的数值仿真表明,能量最优控制器在i.i.d.高斯干扰下表现优于其他控制器。
  • 在高斯随机游走干扰下,路径长度最优控制器能紧密跟踪非因果基准,此时路径长度相对于能量较低。
  • 在对抗性脉冲干扰下,H∞最优和能量最优控制器优于H₂和路径长度最优控制器,验证了其鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。