[论文解读] Hamilton-Jacobi-Bellman Equations for Maximum Entropy Optimal Control
本文提出了一类用于连续时间确定性最优控制的新型哈密顿-雅可比-贝尔曼(HJB)方程,引入了熵正则化,证明了最优值函数是唯一的粘性解。通过广义霍普夫-拉克斯公式建立了计算上的可处理性,并表明对于控制仿射系统,最优控制为高斯分布,在线性二次情况下退化为里卡蒂方程,首次实现了连续时间下的数据驱动、信息论驱动的探索。
Maximum entropy reinforcement learning (RL) methods have been successfully applied to a range of challenging sequential decision-making and control tasks. However, most of existing techniques are designed for discrete-time systems. As a first step toward their extension to continuous-time systems, this paper considers continuous-time deterministic optimal control problems with entropy regularization. Applying the dynamic programming principle, we derive a novel class of Hamilton-Jacobi-Bellman (HJB) equations and prove that the optimal value function of the maximum entropy control problem corresponds to the unique viscosity solution of the HJB equation. Our maximum entropy formulation is shown to enhance the regularity of the viscosity solution and to be asymptotically consistent as the effect of entropy regularization diminishes. A salient feature of the HJB equations is computational tractability. Generalized Hopf-Lax formulas can be used to solve the HJB equations in a tractable grid-free manner without the need for numerically optimizing the Hamiltonian. We further show that the optimal control is uniquely characterized as Gaussian in the case of control affine systems and that, for linear-quadratic problems, the HJB equation is reduced to a Riccati equation, which can be used to obtain an explicit expression of the optimal control. Lastly, we discuss how to extend our results to continuous-time model-free RL by taking an adaptive dynamic programming approach. To our knowledge, the resulting algorithms are the first data-driven control methods that use an information theoretic exploration mechanism in continuous time.
研究动机与目标
- 将最大熵强化学习从离散时间扩展到连续时间确定性最优控制问题。
- 推导一类新的哈密顿-雅可比-贝尔曼(HJB)方程,将熵正则化引入连续时间系统。
- 证明所推导的HJB方程的最优值函数是唯一的粘性解。
- 确保当熵正则化减弱时,渐近一致性成立,极限下恢复标准最优控制。
- 通过广义霍普夫-拉克斯公式实现计算上的可处理性,并为线性二次和控制仿射系统提供显式解。
提出的方法
- 应用动态规划原理,推导具有熵正则化的连续时间确定性系统的HJB方程。
- 采用松弛控制公式,以在确定性系统中建模随机化的控制输入。
- 推导出类似软最大值的哈密顿函数,推广标准哈密顿函数,其结构类似于软Q-learning。
- 采用广义霍普夫-拉克斯公式,以无网格、数值高效的方式求解HJB方程。
- 证明对于控制仿射系统,最优控制被唯一表征为均值等于标准最优控制的高斯分布。
- 在线性二次问题中,将HJB方程简化为代数里卡蒂方程,从而获得显式最优控制表达式。
实验结果
研究问题
- RQ1最大熵原理能否有意义地扩展到连续时间确定性最优控制问题?
- RQ2动态规划原理如何适应于在连续时间下推导具有熵正则化的HJB方程?
- RQ3在连续时间系统中,熵正则化下的最优控制具有何种结构?
- RQ4在最大熵公式下,值函数的正则性如何改善?
- RQ5所得到的HJB方程能否以计算上可处理、无网格的方式求解?
主要发现
- 最大熵控制问题的最优值函数是所推导HJB方程的唯一粘性解。
- 粘性解表现出改进的正则性,其次微分与超微分至多包含一个元素。
- 当温度参数 α → 0 时,值函数一致收敛于标准最优控制问题的值函数,确认了渐近一致性。
- 对于控制仿射系统,最优控制被唯一表征为均值等于标准最优控制的高斯分布。
- 在线性二次问题中,HJB方程退化为代数里卡蒂方程,从而可获得显式解析解。
- 广义霍普夫-拉克斯公式使得HJB方程可在无需哈密顿函数优化的情况下实现无网格数值求解,显著提升了计算可处理性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。