Skip to main content
QUICK REVIEW

[论文解读] Off-Policy Risk-Sensitive Reinforcement Learning Based Constrained Robust Optimal Control

Cong Li, Fangzhou Liu|arXiv (Cornell University)|Jun 10, 2020
Adaptive Dynamic Programming Control参考文献 64被引用 4
一句话总结

该论文提出了一种基于单 critic 的自适应动态规划(ADP)的离策略风险敏感强化学习框架,用于求解在扰动、输入饱和和状态约束下连续时间非线性系统的约束鲁棒最优控制问题。通过将原系统转化为带有风险敏感惩罚项的辅助系统,该方法在确保约束满足和鲁棒稳定性的同时,通过离策略学习和持续激励条件保证了 critic 神经网络权重的收敛性。

ABSTRACT

This paper proposes an off-policy risk-sensitive reinforcement learning based control framework for stabilization of a continuous-time nonlinear system that subjects to additive disturbances, input saturation, and state constraints. By introducing pseudo controls and risk-sensitive input and state penalty terms, the constrained robust stabilization problem of the original system is converted into an equivalent optimal control problem of an auxiliary system. Then, aiming at the transformed optimal control problem, we adopt adaptive dynamic programming (ADP) implemented as a single critic structure to get the approximate solution to the value function of the Hamilton-Jacobi-Bellman (HJB) equation, which results in the approximate optimal control policy that is able to satisfy both input and state constraints under disturbances. By replaying experience data to the off-policy weight update law of the critic artificial neural network, the weight convergence is guaranteed. Moreover, to get experience data to achieve a sufficient excitation required for the weight convergence, online and offline algorithms are developed to serve as principled ways to record informative experience data. The equivalence proof demonstrates that the optimal control strategy of the auxiliary system robustly stabilizes the original system without violating input and state constraints. The proofs of system stability and weight convergence are provided. Simulation results reveal the validity of the proposed control framework.

研究动机与目标

  • 解决传统自适应动态规划(ADP)在具有输入饱和和状态约束的安全关键系统中约束不满足的问题。
  • 开发一种鲁棒控制框架,确保在加性扰动和模型不确定性下仍保持性能与稳定性。
  • 在不依赖演员-评论家相互作用的前提下,确保离策略学习中 critic 神经网络权重的收敛性。
  • 通过单 critic 结构,提供系统稳定性与权重收敛性的严格证明。
  • 通过系统化的在线与离线经验数据采集方法,实现实际应用中的充分激励。

提出的方法

  • 利用伪控制与针对输入和状态约束的风险敏感惩罚项,将原始的约束鲁棒控制问题转化为等价的辅助系统的最优控制问题。
  • 在自适应动态规划(ADP)中采用单 critic 结构,以逼近哈密顿-雅可比-贝尔曼(HJB)方程的解,并直接利用 critic 的值函数推导控制策略。
  • 采用带经验回放的离策略权重更新律,确保 critic 神经网络权重的稳定且收敛的学习过程。
  • 通过在线与离线算法引入持续激励(PE)条件,生成有助于权重收敛的信息性经验数据。
  • 采用基于双曲正切函数的控制策略与风险敏感代价函数,以处理输入饱和,同时保持鲁棒性。
  • 证明辅助系统的最优控制策略可确保原系统在不违反状态或输入约束的前提下实现鲁棒稳定。

实验结果

研究问题

  • RQ1能否设计一种离策略风险敏感强化学习框架,以在加性扰动和输入/状态约束下鲁棒稳定连续时间非线性系统?
  • RQ2在基于 ADP 的控制中,如何在不依赖启发式惩罚函数或变量变换的前提下保证约束满足性?
  • RQ3在离策略学习下,单 critic ADP 结构中 critic 神经网络权重收敛的条件是什么?
  • RQ4在实际中如何实现充分激励,以满足权重收敛所需的持续激励条件?
  • RQ5所提出的方法在不确定性下是否能保持性能与鲁棒性,同时确保严格约束满足?

主要发现

  • 所提出的框架通过将问题转化为辅助系统的等价最优控制问题,确保了在扰动、输入饱和和状态约束下原系统的鲁棒稳定。
  • cnn 权重收敛至一个有界残差集,其大小由系统参数、控制界和近似误差决定。
  • 在满足持续激励条件的前提下,通过经验回放实现的离策略学习可保证权重收敛。
  • 当权重误差超过某一阈值时,系统的李雅普诺夫函数导数为负定,从而确保闭环系统的渐近稳定性。
  • 仿真结果验证了所提方法在扰动下维持约束满足性与鲁棒性能的有效性。
  • 该方法通过在代价函数中显式引入风险敏感性和约束惩罚,实现了性能与安全性的平衡。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。