Skip to main content
QUICK REVIEW

[论文解读] Control Synthesis for Nonlinear Optimal Control via Convex Relaxations

Pengcheng Zhao, Shankar Mohan|arXiv (Cornell University)|Oct 3, 2016
Advanced Control Systems Optimization参考文献 12被引用 4
一句话总结

本文提出一种基于凸松弛的算法,用于求解具有状态和输入约束的非线性最优控制问题的最优控制器。通过将问题表述为在测度上的无限维线性规划,并利用一系列有限维半定规划(SDP)求解,该方法可提取收敛至真实最优控制的多项式控制律,即使最优控制为不连续时亦可实现,同时提供最优代价的可证明收敛下界。

ABSTRACT

This paper addresses the problem of control synthesis for nonlinear optimal control problems in the presence of state and input constraints. The presented approach relies upon transforming the given problem into an infinite-dimensional linear program over the space of measures. To generate approximations to this infinite-dimensional program, a sequence of Semi-Definite Programs (SDP)s is formulated in the instance of polynomial cost and dynamics with semi-algebraic state and bounded input constraints. A method to extract a polynomial control function from each SDP is also given. This paper proves that the controller synthesized from each of these SDPs generates a sequence of values that converge from below to the value of the optimal control of the original optimal control problem. In contrast to existing approaches, the presented method does not assume that the optimal control is continuous while still proving that the sequence of approximations is optimal. Moreover, the sequence of controllers that are synthesized using the presented approach are proven to converge to the true optimal control. The performance of the presented method is demonstrated on three examples.

研究动机与目标

  • 解决具有状态和输入约束的非线性最优控制问题中的控制综合挑战。
  • 开发一种凸优化框架,使控制器能收敛至真实最优控制,且无需假设最优控制的连续性。
  • 提供一种数值可处理的方法,从SDP松弛中生成多项式控制律。
  • 证明由SDP松弛合成的控制器序列收敛至最优控制器。
  • 在双积分器和Dubins汽车等基准非线性系统上验证该方法的有效性。

提出的方法

  • 将最优控制问题重新表述为在非负测度空间上的无限维线性规划(LP),特别采用职业测度。
  • 通过使用矩矩阵和局部化矩阵对多项式动态和代价函数进行松弛,构建一系列有限维半定规划(SDP)的层级。
  • 采用Lasserre的矩-SOS层级方法近似无限维LP,使问题可通过标准SDP求解器数值求解。
  • 开发了一种控制提取过程,从每个SDP松弛的解中恢复多项式反馈控制律。
  • 利用最优控制的弱形式化及职业测度与测试函数之间的对偶性,以保持最优性性质。
  • 在较弱假设下证明了控制器序列收敛至真实最优控制,即使最优控制为不连续时亦成立。

实验结果

研究问题

  • RQ1能否使用凸松弛方法来合成具有状态和输入约束的非线性系统的最优控制器?
  • RQ2所提出的方法是否能保证合成的控制器收敛至真实最优控制,即使最优控制为不连续?
  • RQ3能否从无限维LP公式的SDP松弛中有效提取多项式控制律?
  • RQ4SDP松弛如何同时提供最优代价的下界并实现最优控制的收敛逼近?
  • RQ5能否通过局部近似将该方法应用于具有非多项式动态的系统(如Dubins汽车)?

主要发现

  • 即使最优控制为不连续,由SDP松弛合成的控制器序列仍收敛至真实最优控制,这相比依赖连续性假设的方法具有显著改进。
  • 该方法提供了一组单调收敛至真实最优值的最优代价下界序列。
  • 对于具有自由终端时间的双积分器系统,随着松弛阶数的提高(2k = 6, 8, 12),该方法的控制动作和轨迹与解析计算的最优解高度吻合。
  • 在双积分器的LQR版本中,当2k = 6, 8, 12时,该方法的结果收敛至标准LQR解,表明其在二次代价函数下的准确性。
  • 该方法通过使用动力学的二阶泰勒近似,成功处理了非多项式Dubins汽车模型,得到了可行且收敛的控制律。
  • 控制提取方法成功从SDP解中生成多项式反馈律,使合成控制器可实际部署。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。