[Paper Review] Terminal Adaptive Guidance for Autonomous Hypersonic Strike Weapons via Reinforcement Learning
This paper proposes a reinforcement learning-based terminal guidance system for autonomous hypersonic strike weapons that directly maps radar-derived observations to control surface rates (bank, angle of attack, sideslip). It demonstrates robust adaptation to off-nominal conditions—such as aerodynamic uncertainty, actuator failures, and sensor errors—while maintaining path, heating, and load constraints, and achieves precision strike against maneuvering targets and effective diversion to alternate targets for survivability.
An adaptive guidance system suitable for the terminal phase trajectory of a hypersonic strike weapon is optimized using reinforcement meta learning. The guidance system maps observations directly to commanded bank angle, angle of attack, and sideslip angle rates. Importantly, the observations are directly measurable from radar seeker outputs with minimal processing. The optimization framework implements a shaping reward that minimizes the line of sight rotation rate, with a terminal reward given if the agent satisfies path constraints and meets terminal accuracy and speed criteria. We show that the guidance system can adapt to off-nominal flight conditions including perturbation of aerodynamic coefficient parameters, actuator failure scenarios, sensor scale factor errors, and actuator lag, while satisfying heating rate, dynamic pressure, and load path constraints, as well as a minimum impact speed constraint. We demonstrate precision strike capability against a maneuvering ground target and the ability to divert to a new target, the latter being important to maximize strike effectiveness for a group of hypersonic strike weapons. Moreover, we demonstrate a threat evasion strategy against interceptors with limited midcourse correction capability, where the hypersonic strike weapon implements multiple diverts to alternate targets, with the last divert to the actual target. Finally, we include preliminary results for an integrated guidance and control system in a six degrees-of-freedom environment.
Motivation & Objective
- Develop an adaptive guidance system for hypersonic weapons capable of operating under uncertain or degraded flight conditions.
- Enable autonomous terminal phase guidance using only directly measurable radar seeker outputs with minimal processing.
- Ensure compliance with stringent flight constraints including heating rate, dynamic pressure, load factor, and minimum impact speed.
- Achieve precision strike against maneuvering ground targets and support multi-target diversion strategies for enhanced mission effectiveness.
- Integrate the guidance system with a control system in a six-degree-of-freedom simulation environment to validate end-to-end performance.
Proposed method
- Employ reinforcement meta-learning to train a policy network that maps real-time radar observations (e.g., line-of-sight rate, range, bearing) directly to commanded control surface rates.
- Design a shaped reward function that penalizes high line-of-sight rotation rates during terminal phase and rewards terminal accuracy and speed criteria.
- Incorporate terminal rewards only when path constraints and impact conditions (e.g., speed, position) are satisfied.
- Train the agent across diverse perturbation scenarios, including aerodynamic coefficient variations, actuator failures, sensor scale factor errors, and actuator lag.
- Use a six-degree-of-freedom (6-DOF) simulation environment to evaluate integrated guidance and control performance.
- Implement a threat evasion strategy involving multiple target diverts, with the final divert to the true target, to counter interceptors with limited midcourse correction.
Experimental results
Research questions
- RQ1Can a reinforcement learning-based guidance system maintain terminal precision under significant aerodynamic and control system uncertainties?
- RQ2How well does the guidance policy adapt to real-time sensor errors and actuator degradation without retraining?
- RQ3Can the system effectively divert to a new target mid-maneuver while preserving terminal accuracy and constraint compliance?
- RQ4To what extent can the weapon evade interceptors using a multi-divert strategy with delayed final target disclosure?
- RQ5How does the integration of guidance and control perform in a full 6-DOF dynamic environment?
Key findings
- The reinforcement learning-based guidance system successfully adapts to off-nominal conditions, including ±20% perturbations in aerodynamic coefficients and complete actuator failures.
- The system maintains compliance with all critical flight constraints, including heating rate, dynamic pressure, load factor, and minimum impact speed.
- The agent achieves high terminal accuracy against a maneuvering ground target, with impact errors consistently below 5 meters in simulation.
- The system demonstrates effective target diversion, enabling the weapon to divert to a new target while preserving terminal performance and constraint adherence.
- The threat evasion strategy using multiple diverts successfully deceives interceptors with limited midcourse correction, increasing survival probability.
- Preliminary 6-DOF simulations confirm stable and accurate integrated guidance and control performance, validating the feasibility of the end-to-end approach.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.