[Paper Review] Critic-Only Integral Reinforcement Learning Driven by Variable Gain Gradient Descent for Optimal Tracking Control
This paper proposes a critic-only integral reinforcement learning (IRL) framework with variable gain gradient descent for optimal tracking control of continuous-time nonlinear systems with actuator constraints. By embedding a stabilizing term and adaptive learning rate based on HJB error and Lyapunov derivative, the method eliminates the need for an initial stabilizing controller and dual-approximator networks, achieving uniform ultimate boundedness with a tighter residual set, validated on a 6-DoF UAV model.
Integral reinforcement learning (IRL) was proposed in literature to obviate the requirement of drift dynamics in adaptive dynamic programming framework. Most of the online IRL schemes in literature require two sets of neural network (NNs), known as actor-critic NN and an initial stabilizing controller. Recently, for RL-based robust tracking this requirement of initial stabilizing controller and dual-approximator structure could be obviated by using a modified gradient descent-based update law containing a stabilizing term with critic-only structure. To the best of the authors' knowledge, there has been no study on leveraging such stabilizing term in IRL algorithm framework to solve optimal trajectory tracking problems for continuous time nonlinear systems with actuator constraints. To this end a novel update law leveraging the stabilizing term along with variable gain gradient descent in IRL framework is presented in this paper. With these modifications, the IRL tracking controller can be implemented using only critic NN, while no initial stabilizing controller is required. Another salient feature of the presented update law is its variable learning rate, which scales the pace of learning based on instantaneous Hamilton-Jacobi-Bellman error and rate of variation of Lyapunov function along the system trajectories. The augmented system states and NN weight errors are shown to possess uniform ultimate boundedness (UUB) stability under the presented update law and achieve a tighter residual set. This update law is validated on a full 6-DoF nonlinear model of UAV for attitude control.
Motivation & Objective
- To address the challenge of requiring an initial stabilizing controller in existing online IRL and ADP frameworks for optimal tracking control.
- To eliminate the dual-approximator (actor-critic) structure in IRL-based optimal tracking, reducing computational load.
- To develop a learning rate adaptation mechanism that improves convergence and reduces residual set size in IRL-based optimal control.
- To ensure stability and robustness under actuator constraints and partial system knowledge using only a critic neural network.
- To validate the proposed method on a full 6-DoF nonlinear UAV model with realistic control constraints and dynamics.
Proposed method
- A novel parameter update law is designed using variable gain gradient descent, where the learning rate dynamically adjusts based on the instantaneous Hamilton-Jacobi-Bellman (HJB) error and the rate of change of the Lyapunov function along system trajectories.
- The update law incorporates a stabilizing term within the critic-only structure, enabling policy iteration without requiring an initial stabilizing controller.
- The method operates within an on-policy, online IRL framework, avoiding exploration phases and enabling real-time applicability.
- A single critic neural network approximates the value function, replacing the traditional actor-critic dual-NN architecture and reducing computational complexity.
- The augmented system states and critic weight errors are proven to be uniformly ultimately bounded (UUB) under the proposed update law.
- Persistent excitation is maintained via dithering noise to ensure parameter convergence and avoid rank deficiency in the regressor matrix.
Experimental results
Research questions
- RQ1Can a critic-only IRL framework achieve optimal tracking control for continuous-time nonlinear systems with actuator constraints without requiring an initial stabilizing controller?
- RQ2How does variable gain gradient descent, based on HJB error and Lyapunov derivative, improve convergence and reduce the residual set size compared to constant learning rate methods?
- RQ3What is the stability behavior of the augmented system under the proposed update law, and can uniform ultimate boundedness be guaranteed?
- RQ4To what extent does the critic-only structure reduce computational load compared to dual-approximator actor-critic frameworks in IRL?
- RQ5How well does the proposed method perform in real-world nonlinear systems with partial knowledge of dynamics, such as a 6-DoF UAV?
Key findings
- The proposed critic-only IRL framework with variable gain gradient descent achieves uniform ultimate boundedness (UUB) of the augmented system states and critic weight errors.
- The stabilizing term in the update law successfully eliminates the need for an initial stabilizing controller, a significant advantage over existing IRL and ADP schemes.
- The variable learning rate, based on HJB error and Lyapunov derivative, results in a tighter residual set compared to constant learning rate schemes.
- The critic neural network weights converge to values close to their ideal weights within finite time, as shown in the simulation results.
- The control inputs for elevator, aileron, and rudder remain within the actuator saturation limits of ±90 degrees, confirming feasibility for real-time implementation.
- The value function approximation error approaches near-zero, indicating that the generated control policy is approximately optimal, as confirmed by the near-zero cost in simulation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.