[Paper Review] Dynamic programming for optimal control of stochastic McKean-Vlasov dynamics
This paper establishes a dynamic programming principle for optimal control of stochastic McKean-Vlasov equations with common noise, deriving a Hamilton-Jacobi-Bellman equation via Lions' measure-based differentiability and proving viscosity solution uniqueness. It solves the linear-quadratic control problem explicitly and applies it to a systemic risk model in interbank lending with common noise.
We study the optimal control of general stochastic McKean-Vlasov equation. Such problem is motivated originally from the asymptotic formulation of cooperative equilibrium for a large population of particles (players) in mean-field interaction under common noise. Our first main result is to state a dynamic programming principle for the value function in the Wasserstein space of probability measures, which is proved from a flow property of the conditional law of the controlled state process. Next, by relying on the notion of differentiability with respect to probability measures due to P.L. Lions [32], and It{ô}'s formula along a flow of conditional measures, we derive the dynamic programming Hamilton-Jacobi-Bellman equation, and prove the viscosity property together with a uniqueness result for the value function. Finally, we solve explicitly the linear-quadratic stochastic McKean-Vlasov control problem and give an application to an interbank systemic risk model with common noise.
Motivation & Objective
- To develop a dynamic programming framework for optimal control of stochastic McKean-Vlasov dynamics under common noise.
- To establish a dynamic programming principle in the Wasserstein space of probability measures using flow properties of conditional laws.
- To derive and prove the viscosity solution property and uniqueness of the value function via Lions' calculus and Itô's formula on conditional measures.
- To solve explicitly the linear-quadratic (LQ) stochastic McKean-Vlasov control problem.
- To apply the theoretical results to a systemic risk model in interbank lending with common noise.
Proposed method
- Uses the flow property of the conditional law of the controlled state process to derive the dynamic programming principle in the Wasserstein space.
- Applies Lions' notion of differentiability with respect to probability measures to handle the dependence on the law of the state process.
- Employs Itô’s formula along a flow of conditional measures to derive the Hamilton-Jacobi-Bellman (HJB) equation.
- Proves that the value function is a viscosity solution of the HJB equation and establishes its uniqueness.
- Solves the linear-quadratic (LQ) stochastic McKean-Vlasov control problem by deriving a system of Riccati equations for the coefficients.
- Derives explicit feedback forms for the optimal control and the conditional mean of the state process using the solution of the Riccati system.
Experimental results
Research questions
- RQ1How can a dynamic programming principle be formulated for optimal control of stochastic McKean-Vlasov dynamics in the Wasserstein space?
- RQ2What is the structure of the Hamilton-Jacobi-Bellman equation for such control problems, and how can its viscosity solution property be established?
- RQ3Under what conditions is the value function unique in this class of stochastic control problems?
- RQ4What is the explicit solution to the linear-quadratic (LQ) stochastic McKean-Vlasov control problem with common noise?
- RQ5How can the theoretical framework be applied to model systemic risk in interbank lending with common noise?
Key findings
- The dynamic programming principle is rigorously established in the Wasserstein space of probability measures using the flow property of the conditional law of the controlled state process.
- The value function is proven to be the unique viscosity solution of the derived Hamilton-Jacobi-Bellman equation.
- For the linear-quadratic problem, the Riccati system (5.16) is explicitly solved under the condition $ q^2 \leq \eta $, yielding positive $ \Lambda(t) $, and ensuring existence and uniqueness of solutions for $ \Gamma(t), \gamma(t), \chi(t) $.
- The optimal control is given in feedback form as $ \alpha_t^*(X_t^*) = -(2\Lambda(t)+q)(X_t^* - \mathbb{E}[X_t^*|W^0]) - 2\Gamma(t)\mathbb{E}[X_t^*|W^0] - \gamma(t) $, with explicit time-dependent coefficients.
- The conditional mean $ \bar{X}_t^* = \mathbb{E}[X_t^*|W^0] $ satisfies the SDE $ d\bar{X}_t^* = -(2\Gamma(t)\bar{X}_t^* + \gamma(t))dt + (\sigma_1\bar{X}_t^* + \sigma_0)\rho dW_t^0 $, which reduces to $ \bar{X}_t^* = x_0 + \sigma_0\rho W_t^0 $ when $ \sigma_1 = 0 $.
- The solution recovers and generalizes the result from [18] in the limit of large $ N $, confirming consistency with existing interbank systemic risk models under common noise.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.