Skip to main content
QUICK REVIEW

[Paper Review] Integrated and Adaptive Guidance and Control for Endoatmospheric Missiles via Reinforcement Learning

Brian Gaudet, Roberto Furfaro|arXiv (Cornell University)|Sep 8, 2021
Guidance and Control Systems24 references4 citations
TL;DR

This paper proposes a reinforcement meta-learning framework for integrated, adaptive guidance and control of endoatmospheric missiles, mapping navigation outputs directly to control surface deflections. The system achieves robust intercept performance under large flight envelope variations, off-nominal aerodynamic parameters, flexible body dynamics, and seeker-induced attitude errors, outperforming traditional proportional navigation with a three-loop autopilot in simulation.

ABSTRACT

We apply a reinforcement meta-learning framework to optimize an integrated and adaptive guidance and flight control system for an air-to-air missile. The system is implemented as a policy that maps navigation system outputs directly to commanded rates of change for the missile's control surface deflections. The system induces intercept trajectories against a maneuvering target that satisfy control constraints on fin deflection angles, and path constraints on look angle and load. We test the optimized system in a six degrees-of-freedom simulator that includes a non-linear radome model and a strapdown seeker model, and demonstrate that the system adapts to both a large flight envelope and off-nominal flight conditions including perturbation of aerodynamic coefficient parameters and center of pressure locations, and flexible body dynamics. Moreover, we find that the system is robust to the parasitic attitude loop induced by radome refraction and imperfect seeker stabilization. We compare our system's performance to a longitudinal model of proportional navigation coupled with a three loop autopilot, and find that our system outperforms this benchmark by a large margin. Additional experiments investigate the impact of removing the recurrent layer from the policy and value function networks, performance with an infrared seeker, and flexible body dynamics.

Motivation & Objective

  • To develop an integrated guidance and control system for air-to-air missiles that adapts to large flight envelope variations and off-nominal conditions.
  • To address control constraints on fin deflection, path constraints on look angle and load, and seeker-induced parasitic attitude loops.
  • To replace conventional guidance laws and autopilots with a learned policy that maps navigation outputs directly to control surface commands.
  • To evaluate robustness against aerodynamic coefficient perturbations, center of pressure shifts, and flexible body dynamics.
  • To demonstrate superiority over benchmark proportional navigation with a three-loop autopilot in a high-fidelity six-degree-of-freedom simulation.

Proposed method

  • A deep reinforcement learning policy is trained using a meta-learning framework to generalize across diverse flight conditions.
  • The policy maps real-time navigation system outputs directly to commanded rates of change for missile control surface deflections.
  • The system incorporates recurrent neural networks in both policy and value function networks to model temporal dependencies in control actions.
  • A six-degree-of-freedom simulation environment includes non-linear radome effects and a strapdown seeker model to simulate realistic seeker dynamics.
  • The training environment includes perturbations to aerodynamic coefficients, center of pressure locations, and flexible body dynamics to test robustness.
  • Performance is evaluated against a benchmark longitudinal proportional navigation system with a three-loop autopilot.

Experimental results

Research questions

  • RQ1Can a single reinforcement learning policy effectively handle the full flight envelope of an endoatmospheric missile under varying aerodynamic and dynamic conditions?
  • RQ2How does the inclusion of recurrent layers in the policy and value function networks affect system performance and adaptability?
  • RQ3To what extent is the system robust to seeker-induced parasitic attitude loops caused by radome refraction and imperfect stabilization?
  • RQ4How does the system perform under flexible body dynamics and off-nominal aerodynamic parameter variations?
  • RQ5Does the learned policy outperform classical proportional navigation with a three-loop autopilot in terms of intercept success and constraint satisfaction?

Key findings

  • The reinforcement learning-based system successfully generates intercept trajectories that satisfy both control constraints on fin deflection and path constraints on look angle and load.
  • The system demonstrates robust performance under large flight envelope variations, including perturbations to aerodynamic coefficients and center of pressure locations.
  • The system remains stable and effective despite parasitic attitude loops induced by radome refraction and imperfect seeker stabilization.
  • The inclusion of recurrent layers in the policy and value function networks significantly improves performance, especially in handling temporal dynamics and adapting to changing conditions.
  • The system outperforms the benchmark proportional navigation with a three-loop autopilot in terms of intercept success and constraint adherence.
  • Experiments with an infrared seeker and flexible body dynamics confirm the system's adaptability and robustness across different sensor and airframe configurations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.