[论文解读] Regularity and stability of feedback relaxed controls
本文提出一种带有探索奖励的松弛控制正则化方法,用于构建多维随机退出时间问题的鲁棒反馈控制。它建立了最优反馈控制的霍尔德连续性,并证明了在参数扰动下价值函数与控制的利普希茨稳定性,为基于探索的强化学习启发式方法提供了理论基础。
This paper proposes a relaxed control regularization with general exploration rewards to design robust feedback controls for multi-dimensional continuous-time stochastic exit time problems. We establish that the regularized control problem admits a Hölder continuous feedback control, and demonstrate that both the value function and the feedback control of the regularized control problem are Lipschitz stable with respect to parameter perturbations. Moreover, we show that a pre-computed feedback relaxed control has a robust performance in a perturbed system, and derive a first-order sensitivity equation for both the value function and optimal feedback relaxed control. These stability results provide a theoretical justification for recent reinforcement learning heuristics that including an exploration reward in the optimization objective leads to more robust decision making. We finally prove first-order monotone convergence of the value functions for relaxed control problems with vanishing exploration parameters, which subsequently enables us to construct the pure exploitation strategy of the original control problem based on the feedback relaxed controls.
研究动机与目标
- 解决由于参数扰动和不连续的取最大值映射导致的随机控制问题中反馈控制的不稳定性。
- 开发一种通过探索奖励增强连续时间多维系统鲁棒性的正则化控制框架。
- 在模型不确定性下,建立正则化反馈控制与价值函数的理论稳定性与收敛性性质。
- 通过严格的敏感性与收敛性分析,为基于探索的强化学习启发式方法提供理论依据。
提出的方法
- 引入一种使用一般探索奖励的松弛控制正则化方法,以平滑汉密尔顿-雅可比-贝尔曼(HJB)方程中的取最大值选择。
- 通过分析依赖于探索参数的正则化HJB偏微分方程,建立最优反馈控制的霍尔德连续性。
- 推导出价值函数与最优反馈控制对模型参数扰动的一阶敏感性方程。
- 证明在系数(b, σ, c, f)扰动下,价值函数与反馈控制的利普希茨稳定性。
- 利用插值不等式与霍尔德空间中的先验估计,控制正则化HJB方程解的正则性。
- 证明当探索参数趋于零时,价值函数单调收敛,从而可从松弛控制构造出纯利用策略。
实验结果
研究问题
- RQ1带有探索奖励的正则化反馈控制是否能在多维随机控制问题中实现霍尔德连续性?
- RQ2在模型系数(b, σ, c, f)的小扰动下,价值函数与反馈控制的行为如何?
- RQ3预计算的反馈松弛控制是否能在扰动系统中保持鲁棒性能?
- RQ4价值函数与最优控制对参数变化的一阶敏感性是什么?
- RQ5当探索参数趋于零时,正则化价值函数是否单调收敛至原始控制问题的价值函数?
主要发现
- 正则化问题的最优反馈控制具有霍尔德连续性,确保策略平滑且可实施。
- 在模型系数(b, σ, c, f)扰动下,价值函数与反馈控制具有利普希茨稳定性。
- 预计算的反馈松弛控制在扰动系统中仍能保持鲁棒性能,即使存在微小模型偏差。
- 为价值函数与最优反馈控制推导出一阶敏感性方程,支持对模型不确定性的基于梯度的分析。
- 当探索参数趋于零时,松弛控制问题的价值函数单调收敛至原始控制问题的价值函数。
- 该收敛结果使得可通过消除探索过程,从反馈松弛控制构造出纯利用策略。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。