[Paper Review] Coupling and a generalised Policy Iteration Algorithm in continuous time
This paper presents a generalized policy iteration algorithm for continuous-time controlled diffusion processes with jointly controllable drift and diffusion coefficients. Using mirror coupling of Lindvall and Rogers, it establishes monotonic convergence of payoffs and locally uniform convergence of policies to an optimal solution, proving the algorithm is well-defined and converges to the value function under verifiable conditions on the model data.
We analyse a version of the policy iteration algorithm for the discounted infinite-horizon problem for controlled multidimensional diffusion processes, where both the drift and the diffusion coefficient can be controlled. We prove that, under assumptions on the problem data, the payoffs generated by the algorithm converge monotonically to the value function and an accumulation point of the sequence of policies is an optimal policy. The algorithm is stated and analysed in continuous time and state, with discretisation featuring neither in theorems nor the proofs. A key technical tool used to show that the algorithm is well-defined is the mirror coupling of Lindvall and Rogers.
Motivation & Objective
- To develop a policy iteration algorithm for infinite-horizon discounted control problems in continuous time and state space.
- To prove that the sequence of payoffs generated by the algorithm converges monotonically to the value function.
- To establish that a subsequence of policies converges locally uniformly to an optimal policy.
- To ensure the algorithm is well-defined using classical solutions of the associated PDEs, avoiding reliance on Sobolev spaces.
- To demonstrate convergence using mirror coupling as a key technical tool, even when standard coupling conditions fail.
Proposed method
- The algorithm is formulated in continuous time with no discretization, relying on a generalized policy iteration framework with an arbitrary positive scaling function.
- Mirror coupling from Lindvall and Rogers is used to establish that the payoff function of a locally Lipschitz Markov policy solves the Poisson equation in the classical sense.
- A local path-wise comparison between the time-changed distance process of coupled diffusions and a Bessel process is used to prove high-probability success of the coupling.
- The convergence of payoffs and policies is established through a diagonalization argument and an Arzela-Ascoli-type compactness result.
- The value function is shown to be the pointwise limit of the payoff sequence, and the limiting policy is proven to be optimal.
- Theoretical results are supported by a numerical example showing fast convergence in fewer than six iterations.
Experimental results
Research questions
- RQ1Does the generalized policy iteration algorithm converge monotonically to the value function in continuous-time controlled diffusion processes?
- RQ2Can the policy iteration algorithm be made well-defined using classical solutions of the Hamilton-Jacobi-Bellman equation, rather than generalized solutions in Sobolev spaces?
- RQ3Under what conditions does mirror coupling ensure the payoff function of a locally Lipschitz policy is a classical solution to the Poisson equation?
- RQ4Can the sequence of policies generated by the algorithm converge locally uniformly to an optimal policy?
- RQ5Does the algorithm remain well-defined and convergent when both drift and diffusion coefficients are controlled?
Key findings
- The payoff sequence generated by the policy iteration algorithm converges monotonically to the value function under verifiable assumptions on the model data.
- A subsequence of the policies produced by the algorithm converges locally uniformly to a locally Lipschitz limiting policy.
- The limiting policy is proven to be optimal, with its value function equal to the pointwise limit of the payoff sequence.
- The mirror coupling technique ensures the payoff function of a locally Lipschitz Markov policy is a classical solution to the Poisson equation, even when standard coupling conditions fail.
- The algorithm is well-defined in continuous time with no discretization in theorems or proofs, relying on classical PDE solutions.
- A numerical example demonstrates rapid convergence, finding an optimal policy in fewer than six iterations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.