[Paper Review] Robust Policy Iteration for Continuous-time Linear Quadratic Regulation
This paper establishes the robustness of Kleinman's policy iteration for continuous-time linear quadratic regulation (LQR) under small, bounded disturbances. By modeling policy iteration as a discrete-time nonlinear system, it proves local input-to-state stability, ensuring solutions remain bounded and converge to a small neighborhood of the optimal solution when disturbances are small. The result enables robustness guarantees for off-policy data-driven LQR algorithms under model uncertainty.
This paper studies the robustness of policy iteration in the context of continuous-time infinite-horizon linear quadratic regulation (LQR) problem. It is shown that Kleinman's policy iteration algorithm is inherently robust to small disturbances and enjoys local input-to-state stability in the sense of Sontag. More precisely, whenever the disturbance-induced input term in each iteration is bounded and small, the solutions of the policy iteration algorithm are also bounded and enter a small neighborhood of the optimal solution of the LQR problem. Based on this result, an off-policy data-driven policy iteration algorithm for the LQR problem is shown to be robust when the system dynamics are subjected to small additive unknown bounded disturbances. The theoretical results are validated by a numerical example.
Motivation & Objective
- To analyze the robustness of Kleinman's policy iteration algorithm for continuous-time LQR under small disturbances.
- To establish local input-to-state stability (ISS) of the policy iteration process when disturbances are bounded and small.
- To demonstrate that the algorithm converges to a small neighborhood of the optimal solution even with initial stabilizing control gains and small errors.
- To validate the robustness of off-policy data-driven policy iteration by applying the derived stability results.
Proposed method
- Model the policy iteration process as a discrete-time nonlinear system with disturbances as inputs.
- Prove local exponential stability of the optimal solution in the error-free case using Lyapunov analysis.
- Apply input-to-state stability (ISS) theory to show that bounded, small disturbances result in bounded, small deviations from the optimal solution.
- Use Gronwall's inequality and matrix norm analysis to bound the difference between exact and perturbed policy iteration iterates.
- Leverage the contractive property of the Riccati operator and the invertibility of linear operators in the policy update step.
- Validate the theoretical findings via a numerical example demonstrating convergence under bounded disturbances.
Experimental results
Research questions
- RQ1Is Kleinman’s policy iteration algorithm for continuous-time LQR robust to small, bounded disturbances in the system dynamics or policy evaluation?
- RQ2Does the policy iteration process remain bounded and converge to a small neighborhood of the optimal solution when disturbances are present?
- RQ3Can the theoretical robustness of model-based policy iteration be extended to off-policy data-driven policy iteration algorithms?
- RQ4Under what conditions does the policy iteration algorithm maintain local input-to-state stability in the presence of disturbances?
- RQ5How do errors in system dynamics or function approximation affect the convergence behavior of policy iteration in continuous-time LQR?
Key findings
- The optimal solution of the continuous-time LQR problem is a locally exponentially stable equilibrium of the error-free policy iteration process.
- When disturbances are small and bounded, the policy iteration algorithm is locally input-to-state stable, ensuring bounded and small deviations from the optimal solution.
- For any initial stabilizing control gain, the algorithm converges to a small neighborhood of the optimal solution if disturbances are sufficiently small.
- The off-policy data-driven policy iteration algorithm from [11] is proven robust under small additive, bounded disturbances in the system dynamics.
- Theoretical bounds on the error between exact and perturbed iterates are derived using Gronwall’s inequality and matrix norm analysis.
- Numerical validation confirms that the algorithm remains stable and converges to a small neighborhood of the optimal solution under bounded disturbances.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.