[论文解读] HJB Equations for the Optimal Control of Differential Equations with Delays and State Constraints, II: Optimal Feedbacks and Approximations
本文建立了具有时滞微分方程和状态约束的无限维最优控制问题中最优反馈控制存在的结论,借助基于黏性解的验证定理。证明了在值函数具备适当正则性条件下,由哈密顿-雅可比-贝尔曼方程导出的反馈策略是最优的,并提出了逼近方案,将这些结果推广至点态时滞模型,从而为更广泛的一类问题获得ε-最优控制。
This paper, which is the natural continuation of a previous paper by the same authors, studies a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes some problems arising in economics, in particular the so-called models with time to build. The problem is embedded in a suitable Hilbert space H and the regularity of the associated Hamilton-Jacobi-Bellman (HJB) equation is studied. Therein the main result is that the value function V solves the HJB equation and has continuous classical derivative in the direction of the present. The goal of the present paper is to exploit such result to find optimal feedback strategies for the problem. While it is easy to define formally a feedback strategy in classical sense the proof of its existence and of its optimality is hard due to lack of full regularity of V and to the infinite dimension. Finally, we show some approximation results that allow us to apply our main theorem to obtain epsilon-optimal strategies for a wider class of problems.
研究动机与目标
- 建立由具有状态约束的时滞微分方程所控制的最优控制问题中最优反馈策略的存在性。
- 将验证定理框架扩展至值函数正则性有限的无限维设定。
- 开发逼近技术,使ε-最优反馈策略可应用于未被主定理覆盖的点态时滞问题。
- 弥合理论反馈构造与具有建造时滞动态的经济模型中实际应用之间的差距。
- 在希尔伯特空间设定下,通过基于黏性解的验证定理验证反馈控制的最优性。
提出的方法
- 将时滞控制问题嵌入希尔伯特空间 H = ℝ × L²([-T,0]) 以处理无限维动力学。
- 利用先前研究(Federico et al., 2023)的正则性结果,即值函数在“当前”方向上具有连续的经典导数。
- 应用一种新颖的无限维黏性解验证定理,以确认反馈策略的最优性。
- 从HJB方程形式化构造反馈律,并在充分条件下证明其可容许性与最优性。
- 引入一系列使用光滑化核和截断控制的逼近问题,以处理点态时滞。
- 通过逼近族之间的收敛性论证,表明可为原问题构造出ε-最优策略。
实验结果
研究问题
- RQ1能否在无限维希尔伯特空间中,为具有状态约束的时滞控制问题严格构造并证明最优反馈策略的存在性?
- RQ2当值函数缺乏完整梯度正则性时,如何将验证定理适配至无限维设定?
- RQ3在正则性有限的情况下,何种条件可确保形式化导出的反馈律具有可容许性与最优性?
- RQ4能否通过逼近技术将主定理的适用范围扩展至包含点态时滞的模型(如时间建造经济模型中的模型)?
- RQ5对于未被核心存在性结果覆盖的问题,如何构造ε-最优反馈策略?
主要发现
- 建立了无限维设定下黏性解的验证定理,证明当值函数在“当前”方向上具有连续经典导数时,由HJB方程导出的反馈策略是最优的。
- 提供了充分条件,使得即使缺乏完整梯度正则性,形式化反馈律仍具有可容许性与最优性。
- 通过使用光滑化时滞核和截断控制,构建了处理点态时滞模型的逼近方案。
- 对于任意 ε > 0,通过逼近问题序列,为原始点态时滞问题构造出ε-最优反馈策略。
- 证明了原始问题的值函数是逼近值函数序列的极限,从而确认策略序列的收敛性。
- 当将反馈策略应用于逼近族时,其在原始问题中为3ε-最优,且随着 ε → 0,值函数收敛。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。