[论文解读] Randomization method and backward SDEs for optimal control of partially observed path-dependent stochastic systems
本文通过随机化方法,为部分可观测、路径依赖的随机控制问题的值函数建立了一个倒向SDE表示。通过测度变换重新表述控制问题,并将其与对偶BSDE联系起来,作者建立了一个通用的解决方案框架,该框架超越了马尔可夫性和非退化设定,使数学金融和随机控制中复杂系统得以数值处理。
We consider a unifying framework for stochastic control problem including the following features: partial observation, path-dependence (both with respect to the state and the control), and without any non-degeneracy condition on the stochastic differential equation (SDE) for the controlled state process, driven by a Wiener process. In this context, we develop a general methodology, refereed to as the randomization method, studied in [23] for classical Markovian control under full observation, and consisting basically in replacing the control by an exogenous process independent of the driving noise of the SDE. Our first main result is to prove the equivalence between the primal control problem and the randomized control problem where optimization is performed over change of equivalent probability measures affecting the characteristics of the exogenous process. The randomized problem turns out to be associated by duality and separation argument to a backward SDE, which leads to the so-called randomized dynamic programming principle and randomized equation in terms of the path-dependent filter, and then characterizes the value function of the primal problem. In particular, classical optimal control problems with partial observation affected by non-degenerate Gaussian noise fall within the scope of our framework, and are treated by means of an associated backward SDE.
研究动机与目标
- 为部分可观测、路径依赖的随机控制问题的值函数提供一个通用的后向SDE表示。
- 通过引入路径依赖性和潜变量因子,将现有方法——此前仅限于马尔可夫性或非退化设定——加以扩展。
- 通过测度变换,建立原始部分可观测控制问题与随机化控制问题之间的等价性。
- 通过将随机化公式与对偶后向SDE连接,使问题能够实现数值处理。
- 消除对受控SDE的非退化性条件的依赖,扩大其在现实金融模型中的适用范围。
提出的方法
- 通过Girsanov定理应用参考概率方法,将原始概率测度变换,使得观测过程与受控扩散过程相互独立。
- 通过将控制替换为与驱动噪声独立的外生过程,引入随机化技术,并在影响该过程的等价测度上进行优化。
- 构建一个辅助的随机化控制问题,其中优化对象为测度变换,同时保持原始问题的值不变。
- 构造一个与随机化问题对偶的后向SDE(BSDE),其解可表征原始部分可观测系统值函数。
- 利用非归一化条件分布(通过Zakai方程)表示滤波信号,从而将问题重新表述为全观测控制问题。
- 证明在新概率测度下,随机化跳跃测度的补偿器可确保控制路径度量空间中的收敛性,从而实现逼近。
实验结果
研究问题
- RQ1是否可以在不依赖非退化性或马尔可夫性假设的前提下,为部分可观测、路径依赖的随机控制问题的值函数建立后向SDE表示?
- RQ2随机化方法此前仅用于全观测设定,如何将其适配于具有相关噪声的部分可观测系统?
- RQ3在值等价性和可测控制表示方面,原始控制问题与随机化控制问题之间存在何种关系?
- RQ4所得到的BSDE是否可实现数值逼近?何种条件可确保逼近方案的收敛性?
- RQ5在新概率测度下,随机化跳跃测度的补偿器与原始控制动态之间有何关系?
主要发现
- 原始部分可观测控制问题的值函数由随机化控制公式通过对偶性导出的后向SDE表征。
- 在一般条件下(包括路径依赖系数和相关噪声),原始问题与随机化问题之间的等价性得到严格建立。
- 该方法消除了对扩散系数非退化性的要求,扩展了其在退化模型和潜因子模型中的适用性。
- 随机化控制问题具有BSDE表示,可通过控制路径度量空间中跳跃过程的收敛性实现数值逼近。
- 证明了随机化跳跃测度的补偿器收敛于一个有界且正定的过程,确保了逼近方案的稳定性和收敛性。
- 该框架支持使用Zakai方程进行滤波,并允许即使在路径依赖设定下,值函数也能通过全观测BSDE的解来表示。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。