[论文解读] Exact SDP Formulation for Discrete-Time Covariance Steering with Wasserstein Terminal Cost
该论文为具有Wasserstein距离终端代价的离散时间协方差导向问题提出了一个精确的半定规划(SDP)公式,用随机化状态反馈策略替代了以往的状态历史反馈策略,从而实现了凸的、可计算的优化。其主要贡献是一个SDP,可在无需松弛的情况下获得全局最优的确定性状态反馈控制器,相较于现有方法显著提升了计算效率和性能。
In this paper, we present new results on the covariance steering problem with Wasserstein distance terminal cost. We show that the state history feedback control policy parametrization, which has been used before to solve this class of problems, requires an unnecessarily large number of variables and can be replaced by a randomized state feedback policy which leads to more tractable problem formulations without any performance loss. In particular, we show that under the latter policy, the problem can be equivalently formulated as a semi-definite program (SDP) which is in sharp contrast with our previous results that could only guarantee that the stochastic optimal control problem can be reduced to a difference of convex functions program. Then, we show that the optimal policy that is found by solving the associated SDP corresponds to a deterministic state feedback policy. Finally, we present non-trivial numerical simulations which show the benefits of our proposed randomized state feedback policy derived from the SDP formulation of the problem over existing approaches in the field in terms of computational efficacy and controller performance.
研究动机与目标
- 通过凸优化解决具有Wasserstein终端代价的有限时域协方差导向问题的计算不可行性问题。
- 克服以往基于状态历史反馈策略的局限性,这些策略需要过多的决策变量,导致非凸或松弛化公式。
- 开发一种可计算的、精确的凸公式,避免半定松弛,保证全局最优性。
- 证明从SDP推导出的最优策略是确定性的,简化实现并提升性能。
- 通过数值仿真表明,所提方法在计算速度和控制器性能方面均优于现有方法。
提出的方法
- 引入一种随机化状态反馈策略参数化方法,相比状态历史反馈,显著减少了决策变量数量。
- 通过变量变换,将具有Wasserstein距离终端代价的协方差导向问题公式化为凸的半定规划(SDP)。
- 以一种使SDP目标函数关于决策变量呈线性形式的方式表达Wasserstein距离,从而实现高效求解。
- 证明SDP的最优解对应于一个确定性状态反馈策略,消除了控制律中的随机性。
- 利用策略参数化中的块对角结构,在保持最优性的同时降低计算复杂度。
- 在数值仿真中,将基于SDP的策略与截断和非截断状态历史反馈以及仿射扰动反馈策略进行对比实现与比较。
实验结果
研究问题
- RQ1具有Wasserstein终端代价的协方差导向问题能否被公式化为无需松弛的精确凸优化问题?
- RQ2相较于状态历史反馈,随机化状态反馈策略是否在更少决策变量下实现相当或更优的性能?
- RQ3从SDP公式推导出的最优策略是否为确定性?这对控制器设计有何影响?
- RQ4与现有非凸或松弛化方法相比,所提SDP方法的计算复杂度随问题时域长度的扩展如何变化?
- RQ5所提SDP公式是否在解的质量和计算时间方面均优于现有策略参数化方法?
主要发现
- 所提出的SDP公式是精确的,不依赖任何凸松弛,可保证全局最优性。
- 尽管采用了随机化策略参数化,从SDP推导出的最优策略是确定性的,简化了实现。
- 决策变量数量随问题时域长度$N$线性增长,计算复杂度为$\tfrac{1}{3}$阶,如表I所示。
- SDP的计算时间随$\mathcal{O}(N^3)$增长,2D双积分器系统中$N=150$时报告的计算时间为29.34秒。
- 数值仿真表明,基于SDP的状态反馈策略在性能和计算时间方面均优于截断状态历史反馈和仿射扰动反馈策略。
- 随着$N$增大,基于非截断历史的策略计算时间变得不可行,而基于SDP的方法仍保持可扩展性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。