Skip to main content
QUICK REVIEW

[论文解读] Regret-optimal Estimation and Control

Gautam Goel, Babak Hassibi|arXiv (Cornell University)|Jun 22, 2021
Advanced Control Systems Optimization被引用 4
一句话总结

该论文通过最小化在线因果策略与全知非因果基准之间的最坏情况差异(即遗憾),为线性时变系统引入了遗憾最优的估计器和控制器。利用算子理论技术和状态空间重构方法,推导出因果滤波器和控制器,实现了以扰动能量为度量的紧致、数据相关的遗憾边界,其性能在数值实验中优于标准方法如EKF和MPC。

ABSTRACT

We consider estimation and control in linear time-varying dynamical systems from the perspective of regret minimization. Unlike most prior work in this area, we focus on the problem of designing causal estimators and controllers which compete against a clairvoyant noncausal policy, instead of the best policy selected in hindsight from some fixed parametric class. We show that the regret-optimal estimator and regret-optimal controller can be derived in state-space form using operator-theoretic techniques from robust control and present tight,data-dependent bounds on the regret incurred by our algorithms in terms of the energy of the disturbances. Our results can be viewed as extending traditional robust estimation and control, which focuses on minimizing worst-case cost, to minimizing worst-case regret. We propose regret-optimal analogs of Model-Predictive Control (MPC) and the Extended KalmanFilter (EKF) for systems with nonlinear dynamics and present numerical experiments which show that our regret-optimal algorithms can significantly outperform standard approaches to estimation and control.

研究动机与目标

  • 解决传统H₂和H∞控制的局限性,后者假设特定的扰动类别,可能在扰动不匹配时失效。
  • 设计自适应的因果估计器与控制器,使其在与具备未来扰动全知能力的全局最优非因果基准对比时,最小化遗憾。
  • 将遗憾最小化概念从固定参数策略类扩展到更一般的非参数比较,与最优全知策略进行对比。
  • 为估计与控制问题提供以扰动能量为度量的紧致、数据相关的遗憾边界。

提出的方法

  • 将遗憾最小化表述为因果在线策略与已知全序列扰动的全知非因果策略之间的比较。
  • 利用鲁棒控制中的算子理论技术,推导出遗憾最优估计器与控制器的状态空间表示。
  • 构建一个包含3n个状态的扩展系统,其中新系统中的H∞滤波器对应于原系统中的遗憾最优滤波器。
  • 通过状态扩展方法,将结果推广至具有预测时域和控制时延的场景,重新表述系统动力学。
  • 利用相同框架推导出非线性系统的遗憾最优模型预测控制(MPC)与扩展卡尔曼滤波器(EKF)类比。
  • 建立紧致的、基于能量的遗憾边界,其大小与总扰动能量成正比,且与扰动结构无关。

实验结果

研究问题

  • RQ1能否设计一个因果估计器,使其在与可提前访问所有测量值的平滑非因果估计器对比时,最小化遗憾?
  • RQ2能否设计一个因果控制器,使其在与已知所有未来扰动的全知控制器对比时,最小化遗憾?
  • RQ3如何以一种避免对策略类施加严格参数假设的方式,形式化遗憾最优的估计与控制?
  • RQ4对于线性时变系统,以扰动能量为度量的遗憾的紧致、数据相关边界是什么?
  • RQ5如何将遗憾最优框架扩展至具有预测或控制时延的系统?

主要发现

  • 遗憾最优滤波器被表示为扩展3n状态系统中的H∞滤波器,可作为标准滤波器(如卡尔曼滤波器和H∞滤波器)的即插即用替代方案。
  • 遗憾最优控制器的最坏情况遗憾被一个与扰动能量成正比的常数所界定,且与扰动分布无关。
  • 数值实验表明,遗憾最优算法在估计与控制代价方面显著优于标准EKF和MPC。
  • 通过状态扩展,该框架可推广至具有预测时域和控制时延的系统,同时保持遗憾最优性。
  • 遗憾边界是紧致且数据相关的,与总扰动能量呈线性关系,使性能保证具有鲁棒性与可解释性。
  • 该方法通过遗憾最优的EKF与MPC类比扩展至非线性系统,展示了其在非线性模型之外的实际适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。