[论文解读] A uniform Tauberian theorem in optimal control
本文在连续时间最优控制中建立了统一的Tauberian定理,证明了当 $T \to \infty$ 时,有限时域值函数 $V_T(x)$ 的一致收敛性等价于当 $\lambda \to 0$ 时贴现值函数 $V_\lambda(x)$ 的一致收敛性,且极限相同。该结果将经典的Hardy-Littlewood与Feller Tauberian定理推广至无需遍历性假设的受控动力系统,通过离散时间逼近与等价性论证实现。
In an optimal control framework, we consider the value $V_T(x)$ of the problem starting from state $x$ with finite horizon $T$, as well as the value $V_λ(x)$ of the $λ$-discounted problem starting from $x$. We prove that uniform convergence (on the set of states) of the values $V_T(\cdot)$ as $T$ tends to infinity is equivalent to uniform convergence of the values $V_λ(\cdot)$ as $λ$ tends to 0, and that the limits are identical. An example is also provided to show that the result does not hold for pointwise convergence. This work is an extension, using similar techniques, of a related result in a discrete-time framework \cite{LehSys}.
研究动机与目标
- 在无需遍历性假设的前提下,为连续时间最优控制问题建立统一的Tauberian定理。
- 解决一个开放问题:有限时域值函数 $V_T(x)$ 的一致收敛性是否蕴含贴现值函数 $V_\lambda(x)$ 的一致收敛性,反之亦然。
- 证明在一致收敛条件下,$V_T(x)$ 与 $V_\lambda(x)$ 的极限完全一致。
- 将Lehrer与Sorin(2012)在离散时间下的结果推广至连续时间受控动力系统。
提出的方法
- 构造一个状态空间为 $\widetilde{\Omega} = \Omega \times [0,1]$ 且代价函数为 $\widetilde{g}(\omega,x) = x$ 的离散时间动态规划问题。
- 定义一个多值转移映射 $\widetilde{\Gamma}$,用于编码连续时间系统在单位时间间隔内的轨迹。
- 应用Lehrer与Sorin(2012)关于离散时间统一Tauberian收敛性的主要结果于所构造的离散系统。
- 建立连续时间值函数 $V_t(\omega)$ 与离散时间值函数 $v_n(\omega,x)$ 之间的统一有界性,证明 $|V_t(\omega) - v_{\lfloor t \rfloor}(\omega,x)| \leq \frac{2}{\lfloor t \rfloor}$。
- 利用积分不等式 $\lambda \int_0^\infty |(1-\lambda)^{\lfloor t \rfloor} - e^{-\lambda t}| dt \to 0$,证明当 $\lambda \to 0$ 时,$|V_\lambda(\omega) - v_\lambda(\omega,x)|$ 一致收敛于零。
- 借助离散时间中的收敛性等价性,推导出连续时间下的统一等价性。
实验结果
研究问题
- RQ1当 $T \to \infty$ 时,有限时域值函数 $V_T(x)$ 的一致收敛性是否蕴含当 $\lambda \to 0$ 时贴现值函数 $V_\lambda(x)$ 的一致收敛性?
- RQ2反之是否成立:$V_\lambda(x)$ 的一致收敛性是否蕴含 $V_T(x)$ 的一致收敛性?
- RQ3当两者均一致收敛时,$V_T(x)$ 与 $V_\lambda(x)$ 的极限是否相同?
- RQ4经典的Hardy-Littlewood与Feller Tauberian定理能否推广至无需遍历性的受控连续时间系统?
- RQ5有限时域与贴现收敛性之间的等价性在非遍历动力系统下是否依然稳健?
主要发现
- 当 $T \to \infty$ 时,$V_T(\cdot)$ 的一致收敛性与当 $\lambda \to 0$ 时,$V_\lambda(\cdot)$ 的一致收敛性等价。
- 当两者均一致收敛时,$V_T(\cdot)$ 与 $V_\lambda(\cdot)$ 的极限完全相同。
- 该等价性不依赖于系统的遍历性或可控制性假设。
- 由于有界性 $\lambda \int_0^\infty |(1-\lambda)^{\lfloor t \rfloor} - e^{-\lambda t}| dt \to 0$,$|V_\lambda(\omega) - v_\lambda(\omega,x)|$ 在 $\omega$ 与 $x$ 上关于 $\lambda \to 0$ 一致收敛于零。
- 本结果将Lehrer与Sorin(2012)的离散时间统一Tauberian定理推广至连续时间最优控制。
- 提供了一个反例,表明该等价性在逐点收敛下不成立,凸显了一致收敛的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。