Skip to main content
QUICK REVIEW

[论文解读] Duality between density function and value function with applications in constrained optimal control and Markov Decision Process

Yuxiao Chen, Aaron D. Ames|arXiv (Cornell University)|Feb 25, 2019
Traffic control and management参考文献 21被引用 7
一句话总结

本文在最优控制与马尔可夫决策过程(MDPs)中建立了价值函数与密度函数之间的对偶性,使得诸如安全性和容量约束等约束最优控制问题可被重新表述为对密度函数的优化问题。通过采用一种原始-对偶算法,该算法在求解标准HJB偏微分方程(PDE)以获得价值函数与通过Liouville方程演化密度函数之间交替进行,该方法在保持最优性的同时高效地实现了约束,已在机器人导航、交通控制以及受扰动影响下的实验性Segway控制中取得成功。

ABSTRACT

Density function describes the density of states in the state space of a dynamic system or a Markov Decision Process (MDP). Its evolution follows the Liouville equation. We show that the density function is the dual of the value function in the optimal control problems. By utilizing the duality, constraints that are hard to enforce in the primal value function optimization such as safety constraints in robot navigation, traffic capacity constraints in traffic flow control can be posed on the density function, and the constrained optimal control problem can be solved with a primal-dual algorithm that alternates between the primal and dual optimization. The primal optimization follows the standard optimal control algorithm with a perturbation term generated by the density constraint, and the dual problem solves the Liouville equation to get the density function under a fixed control strategy and updates the perturbation term. Moreover, the proposed method can be extended to the case with exogenous disturbance, and guarantee robust safety under the worst-case disturbance. We apply the proposed method to three examples, a robot navigation problem and a traffic control problem in sim, and a segway control problem with experiment.

研究动机与目标

  • 建立最优控制与MDPs中价值函数与密度函数之间的理论对偶性。
  • 解决在原始价值函数形式下难以直接强制执行的约束最优控制问题,如安全性与容量约束。
  • 开发一种原始-对偶算法,该算法在价值函数优化(HJB PDE)与密度函数演化(Liouville方程)之间交替进行,同时满足约束。
  • 将该框架扩展至具有外生扰动的系统以及多个耦合的MDPs。
  • 通过机器人导航、交通控制和Segway稳定性的仿真与实验结果验证该方法。

提出的方法

  • 提出一种对偶框架,使密度函数在连续时间系统与离散MDPs中均成为价值函数的对偶。
  • 采用基于常微分方程(ODE)的方法,通过在固定控制策略下正向求解Liouville方程来计算密度函数。
  • 引入一种原始-对偶算法:原始步骤通过引入源自密度约束的扰动项来求解HJB PDE,对偶步骤则通过Liouville方程演化密度函数。
  • 采用定点迭代方法,根据当前密度更新扰动项,以确保约束被满足。
  • 通过将最坏情况扰动纳入密度演化过程,将方法扩展至鲁棒控制,从而确保鲁棒安全性。
  • 在仅有模拟器可用时,采用核密度估计来近似密度函数,使该方法可与强化学习集成。

实验结果

研究问题

  • RQ1在连续最优控制与离散MDPs中,密度函数能否被形式化地确立为价值函数的对偶?
  • RQ2如何通过在密度函数上进行优化,自然地实现如安全性和交通容量等约束?
  • RQ3能否设计一种原始-对偶算法,使其在优化价值函数与演化密度函数之间交替进行,以满足约束?
  • RQ4该方法在外部扰动下表现如何?能否保证鲁棒安全性?
  • RQ5该框架能否扩展至多个耦合的MDPs,并与强化学习集成?

主要发现

  • 在连续系统与MDPs中,密度函数被证明是价值函数的对偶,其对偶性根植于Liouville方程与HJB方程。
  • 所提出的原始-对偶算法成功地强制执行了密度约束,例如将第7号区域的交通密度限制在指定范围内,受限解的代价为79.21,而无约束情况下的代价为71.05。
  • 在受限MDP情形下,最优策略变为随机策略,瓶颈区域的密度恰好达到其允许的最大值,证实了约束的满足。
  • 通过将扰动动力学纳入Liouville方程以实现密度演化,该方法实现了对最坏情况扰动的鲁棒安全性。
  • 该框架已在Segway系统上通过实验验证,展示了其在真实世界中的适用性与在控制约束下的稳定性。
  • 该方法可通过核密度估计扩展至强化学习,使基于密度的安全约束能够在基于模拟的学习中实现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。