Skip to main content
QUICK REVIEW

[论文解读] Dynamic Programming Principles for Optimal Stopping with Expectation Constraint

Erhan Bayraktar, Song Yao|arXiv (Cornell University)|Aug 7, 2017
Stochastic processes and financial applications参考文献 39被引用 3
一句话总结

本文为累积成本具有期望约束的最优停止问题建立了动态规划原理(DPP),证明了价值函数的连续性,并将其表征为一个完全非线性抛物型汉密尔顿-雅可比-贝尔曼方程的粘性解。其关键创新在于引入条件期望成本作为额外的状态过程,从而实现了对约束最优停止问题中一个开放问题的解决。

ABSTRACT

We analyze an optimal stopping problem with a constraint on the expected cost. When the reward function and cost function are Lipschitz continuous in state variable, we show that the value of such an optimal stopping problem is a continuous function in current state and in budget level. Then we derive a dynamic programming principle (DPP) for the value function in which the conditional expected cost acts as an additional state process. As the optimal stopping problem with expectation constraint can be transformed to a stochastic optimization problem with supermartingale controls, we explore a second DPP of the value function and thus resolve an open question recently raised in [S. Ankirchner, M. Klein, and T. Kruse, A verification theorem for optimal stopping problems with expectation constraints, Appl. Math. Optim., 2017, pp. 1-33]. Based on these two DPPs, we characterize the value function as a viscosity solution to the related fully non-linear parabolic Hamilton-Jacobi-Bellman equation.

研究动机与目标

  • 为解决累积成本具有期望约束的最优停止问题中缺乏动态规划原理的问题。
  • 在 Lipschitz 连续性和非退化条件下,证明价值函数在状态变量和预算变量上的连续性。
  • 通过引入条件期望成本作为辅助状态过程,解决约束最优停止问题中DPP的开放问题。
  • 将价值函数表征为完全非线性抛物型 HJB 方程的粘性解。
  • 将约束最优停止问题与具有上鞅控制的随机控制问题联系起来,从而推导出第二个 DPP。

提出的方法

  • 通过将条件期望成本视为额外的状态过程,利用平移的随机微分方程和流性质推导出 DPP。
  • 采用先验估计和近似停止策略,证明价值函数在 (t, x, y) 上的连续性。
  • 利用正规条件概率分布和状态过程的马氏性质,建立 DPP 的 '≤' 方向。
  • 通过粘贴局部停止规则构造 ε-最优策略,利用价值函数的连续性证明 DPP 的 '≥' 方向。
  • 将约束最优停止问题转化为从预算水平 y 出发的具有上鞅控制的无约束随机控制问题。
  • 将价值函数表征为一个完全非线性抛物型 HJB 方程的粘性解,其中包含一种 Monge-Ampère 类型的结构。

实验结果

研究问题

  • RQ1能否为累积成本具有期望约束的最优停止问题建立动态规划原理?
  • RQ2在 Lipschitz 连续性和非退化条件下,价值函数在状态变量和预算水平上是否连续?
  • RQ3如何将条件期望成本作为状态过程引入,以在该约束设定下实现 DPP?
  • RQ4能否将约束问题重新表述为具有上鞅控制的随机控制问题,从而推导出第二个 DPP?
  • RQ5价值函数是否为完全非线性抛物型汉密尔顿-雅可比-贝尔曼方程的粘性解?

主要发现

  • 在 f、π、g 满足 Lipschitz 连续性且 g 非退化时,价值函数 V(t, x, y) 在 (t, x, y) 上连续。
  • 建立了动态规划原理,其中条件期望成本 Y^{t,x,τ}_s 作为额外的状态过程,得到如下形式的 DPP:V(t,x,y) = sup_E[1_{τ≤ζ(τ)} R(t,x,τ) + 1_{τ>ζ(τ)} (V(ζ(τ), X^{t,x}_{ζ(τ)}, Y^{t,x,τ}_{ζ(τ)}) + ∫_t^{ζ(τ)} f(r,X^{t,x}_r) dr)]。
  • 价值函数被表征为一个完全非线性抛物型汉密尔顿-雅可比-贝尔曼方程的粘性解,包含一种 Monge-Ampère 类型的结构。
  • 通过将约束问题转化为具有上鞅控制的无约束随机控制问题,推导出第二个 DPP,解决了 [Ankirchner et al., 2017] 中的开放问题。
  • 价值函数的连续性使得可通过粘贴法构造 ε-最优停止策略,这对证明 DPP 的反向不等式至关重要。
  • 证明了价值函数是有限且可测的,且满足 E_t[K_*] < ∞,从而确保了动态规划框架的有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。