Skip to main content
QUICK REVIEW

[论文解读] Policy Gradient-based Algorithms for Continuous-time Linear Quadratic Control

Jingjing Bu, Afshin Mesbahi|arXiv (Cornell University)|Jun 12, 2020
Adaptive Dynamic Programming Control参考文献 29被引用 11
一句话总结

本文提出基于策略梯度的算法以解决连续时间线性二次调节器(LQR)问题,将最优控制增益表述为稳定反馈增益上的矩阵函数。该工作建立了光滑性、强制性及梯度主导性等性质,从而为梯度流、自然梯度流及拟牛顿流提供收敛性保证,在适当的步长规则下实现线性和Q-二次收敛速率,并通过投影梯度下降方法扩展至稀疏控制合成,实现次线性收敛。

ABSTRACT

We consider the continuous-time Linear-Quadratic-Regulator (LQR) problem in terms of optimizing a real-valued matrix function over the set of feedback gains. The results developed are in parallel to those in Bu et al. [1] for discrete-time LTI systems. In this direction, we characterize several analytical properties (smoothness, coerciveness, quadratic growth) that are crucial in the analysis of gradient-based algorithms. We also point out similarities and distinctive features of the continuous time setup in comparison with its discrete time analogue. First, we examine three types of well-posed flows direct policy update for LQR: gradient flow, natural gradient flow and the quasi-Newton flow. The coercive property of the corresponding cost function suggests that these flows admit unique solutions while the gradient dominated property indicates that the underling Lyapunov functionals decay at an exponential rate; quadratic growth on the other hand guarantees that the trajectories of these flows are exponentially stable in the sense of Lyapunov. We then discuss the forward Euler discretization of these flows, realized as gradient descent, natural gradient descent and quasi-Newton iteration. We present stepsize criteria for gradient descent and natural gradient descent, guaranteeing that both algorithms converge linearly to the global optima. An optimal stepsize for the quasi-Newton iteration is also proposed, guaranteeing a $Q$-quadratic convergence rate--and in the meantime--recovering the Kleinman-Newton iteration. Lastly, we examine LQR state feedback synthesis with a sparsity pattern. In this case, we develop the necessary formalism and insights for projected gradient descent, allowing us to guarantee a sublinear rate of convergence to a first-order stationary point.

研究动机与目标

  • 开发适用于连续时间LQR的基于梯度的优化方法,直接更新反馈增益,无需求解Riccati方程。
  • 建立LQR代价函数在稳定反馈增益上的解析性质——光滑性、强制性、梯度主导性及二次增长性。
  • 分析连续时间流(梯度、自然梯度、拟牛顿)及其收敛性,基于李雅普诺夫稳定性与李雅普诺夫泛函的指数衰减。
  • 推导离散时间对应方法(梯度下降、自然梯度下降、拟牛顿)的步长准则,确保线性与Q-二次收敛。
  • 将框架扩展至具有稀疏模式的结构化LQR问题,采用投影梯度下降,保证收敛至一阶平稳点。

提出的方法

  • 通过在一组线性无关初始条件下取平均,将无限时域LQR代价重新表述为稳定反馈增益上的矩阵函数,消除对初始状态的依赖。
  • 证明所得代价函数具有光滑性、强制性及梯度主导性,确保紧致的子水平集与全局最小值的存在性。
  • 分析三种连续时间流:梯度流、自然梯度流与拟牛顿流,表明李雅普诺夫泛函指数衰减,轨迹在李雅普诺夫意义下指数稳定。
  • 推导上述流的前向欧拉离散化形式,对应梯度下降、自然梯度下降与拟牛顿迭代,并给出确保线性收敛的显式步长条件。
  • 为拟牛顿迭代提出最优步长,实现Q-二次收敛,并恢复经典的Kleinman-Newton方法。
  • 为具有稀疏性约束的结构化LQR问题,构建投影梯度下降框架,通过将更新投影至稀疏模式以保持结构保真度。

实验结果

研究问题

  • RQ1能否在不求解代数Riccati方程的前提下,严格地将策略梯度方法应用于连续时间LQR问题?
  • RQ2在连续时间下,LQR代价函数在反馈增益上的解析性质(光滑性、强制性、梯度主导性)为何?
  • RQ3连续时间策略更新流(梯度、自然梯度、拟牛顿)是否具有唯一解,并表现出李雅普诺夫泛函的指数衰减与轨迹的指数稳定性?
  • RQ4何种步长条件可保证离散时间梯度下降与自然梯度下降在连续时间LQR中的线性收敛?
  • RQ5能否使用投影梯度下降求解具有稀疏性约束的结构化LQR问题?可保证何种收敛速率?

主要发现

  • 在稳定反馈增益上的LQR代价函数具有光滑性、强制性与梯度主导性,确保子水平集紧致且全局最小值存在。
  • 梯度流、自然梯度流与拟牛顿流均具有唯一解,且李雅普诺夫泛函指数衰减,轨迹在李雅普诺夫意义下指数稳定。
  • 在推导的步长准则下,梯度下降与自然梯度下降线性收敛至全局最优解,且步长在子水平集上远离零。
  • 为拟牛顿迭代设计的最优步长可实现Q-二次收敛,并恢复经典Kleinman-Newton迭代方法。
  • 针对具有稀疏模式的结构化LQR问题,投影梯度下降可实现收敛至一阶平稳点的次线性速率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。