Skip to main content
QUICK REVIEW

[论文解读] Adaptive Regret for Control of Time-Varying Dynamics

Paula Gradu, Elad Hazan|arXiv (Cornell University)|Jul 8, 2020
Advanced Bandit Algorithms Research参考文献 31被引用 14
一句话总结

本文提出了一种新颖的元算法,可将任意具有次线性遗憾的控制器转换为适用于时变线性动态系统的次线性自适应遗憾控制器。通过利用带记忆的在线凸优化及一种新颖的自适应遗憾界,该方法使控制器能够动态适应变化的系统动态,从而在非随机环境中实现近乎紧致的遗憾界。

ABSTRACT

We consider the problem of online control of systems with time-varying linear dynamics. This is a general formulation that is motivated by the use of local linearization in control of nonlinear dynamical systems. To state meaningful guarantees over changing environments, we introduce the metric of {\it adaptive regret} to the field of control. This metric, originally studied in online learning, measures performance in terms of regret against the best policy in hindsight on {\it any interval in time}, and thus captures the adaptation of the controller to changing dynamics. Our main contribution is a novel efficient meta-algorithm: it converts a controller with sublinear regret bounds into one with sublinear {\it adaptive regret} bounds in the setting of time-varying linear dynamical systems. The main technical innovation is the first adaptive regret bound for the more general framework of online convex optimization with memory. Furthermore, we give a lower bound showing that our attained adaptive regret bound is nearly tight for this general framework.

研究动机与目标

  • 为通过局部线性化控制具有时变动态的非线性动态系统而提出解决方案。
  • 为在标准遗憾度量失效的非平稳环境中实现在线控制提供性能保证。
  • 在控制理论中引入自适应遗憾作为有意义的性能度量,以捕捉适应变化动态的能力。
  • 设计一种高效的元算法,将标准的遗憾最小化控制器转换为自适应遗憾最小化控制器。
  • 为带记忆的在线凸优化建立近乎紧致的自适应遗憾下界。

提出的方法

  • 提出一种元算法,通过在代理损失函数上使用在线梯度下降来维护和更新多个专家控制器。
  • 使用分层集合系统 $S_t$ 来维护和更新控制器权重,确保时间复杂度为对数级。
  • 采用代理代理代价函数来估计未来性能,并在每个时间步指导控制器选择。
  • 将该框架应用于带记忆的在线凸优化,推导出该一般设置下的首个自适应遗憾界。
  • 提出一种新颖的权重更新规则,可在任意时间区间内动态将控制转移至表现最佳的专家。
  • 使用自动微分在每一步计算时变线性化系统动态 $(A_t, B_t)$。

实验结果

研究问题

  • RQ1我们能否设计一种控制器,在非随机环境中有效适应时变线性动态?
  • RQ2如何将在线控制中的遗憾最小化扩展以捕捉对系统行为变化的适应能力?
  • RQ3在时变系统中,带记忆的在线凸优化的最优可实现自适应遗憾界是什么?
  • RQ4元算法能否高效地将标准遗憾最小化控制器转换为具有次线性自适应遗憾的控制器?
  • RQ5所推导的自适应遗憾界是否在对数因子范围内为紧致?

主要发现

  • 所提出的元算法在时变线性动态系统中实现了次线性自适应遗憾,从而能够有效适应变化的动态。
  • 该方法首次为带记忆的在线凸优化建立了自适应遗憾界,该通用框架具有广泛的影响。
  • 证明了近乎紧致的下界,表明所获得的自适应遗憾界在对数因子范围内为最优。
  • 实验表明,MARC元算法在具有时变动态的非平稳环境中优于基线控制器(如GPC和iLQR)。
  • 该算法每时间步维持 $O(\log T)$ 个工作集合,确保在动态控制器选择下仍具有计算效率。
  • 超参数调优表明,该方法具有鲁棒性,实验中在 $\eta = 0.05$ 和 $\sigma = 10^{-2}$ 时达到最优性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。