Skip to main content
QUICK REVIEW

[论文解读] Terminal Adaptive Guidance for Autonomous Hypersonic Strike Weapons via Reinforcement Learning

Brian Gaudet, Roberto Furfaro|arXiv (Cornell University)|Oct 1, 2021
Guidance and Control Systems参考文献 25被引用 4
一句话总结

本文提出了一种基于强化学习的自主高超声速打击武器终端制导系统,可直接将雷达获取的观测数据映射为控制面速率(滚转、迎角、侧滑)。该系统在面对非 nominal 条件(如气动不确定性、执行器故障和传感器误差)时表现出鲁棒性,同时保持路径、热负荷和载荷约束,能够精确打击机动目标,并有效实现对备选目标的偏转以提升生存能力。

ABSTRACT

An adaptive guidance system suitable for the terminal phase trajectory of a hypersonic strike weapon is optimized using reinforcement meta learning. The guidance system maps observations directly to commanded bank angle, angle of attack, and sideslip angle rates. Importantly, the observations are directly measurable from radar seeker outputs with minimal processing. The optimization framework implements a shaping reward that minimizes the line of sight rotation rate, with a terminal reward given if the agent satisfies path constraints and meets terminal accuracy and speed criteria. We show that the guidance system can adapt to off-nominal flight conditions including perturbation of aerodynamic coefficient parameters, actuator failure scenarios, sensor scale factor errors, and actuator lag, while satisfying heating rate, dynamic pressure, and load path constraints, as well as a minimum impact speed constraint. We demonstrate precision strike capability against a maneuvering ground target and the ability to divert to a new target, the latter being important to maximize strike effectiveness for a group of hypersonic strike weapons. Moreover, we demonstrate a threat evasion strategy against interceptors with limited midcourse correction capability, where the hypersonic strike weapon implements multiple diverts to alternate targets, with the last divert to the actual target. Finally, we include preliminary results for an integrated guidance and control system in a six degrees-of-freedom environment.

研究动机与目标

  • 开发一种自适应制导系统,使高超声速武器能够在存在不确定或退化飞行条件的情况下运行。
  • 仅利用可直接测量的雷达导引头输出(经最少处理)实现自主终端阶段制导。
  • 确保符合严格的飞行约束条件,包括热流率、动压、载荷因子和最小撞击速度。
  • 实现对机动地面目标的精确打击,并支持多目标偏转策略以提升任务有效性。
  • 在六自由度(6-DOF)仿真环境中将制导系统与控制系统集成,以验证端到端性能。

提出的方法

  • 采用强化元学习训练策略网络,将实时雷达观测(如视线角速率、距离、方位角)直接映射为指令控制面速率。
  • 设计形状奖励函数,惩罚终端阶段过高的视线旋转速率,同时奖励终端精度和速度标准。
  • 仅在路径约束和撞击条件(如速度、位置)满足时才施加终端奖励。
  • 在多种扰动场景下训练智能体,包括气动系数变化、执行器故障、传感器比例因子误差和执行器延迟。
  • 使用六自由度(6-DOF)仿真环境评估集成制导与控制性能。
  • 实施一种威胁规避策略,包含多次目标偏转,最终偏转至真实目标,以应对具备有限中段修正能力的拦截器。

实验结果

研究问题

  • RQ1基于强化学习的制导系统是否能在显著的气动和控制系统不确定性下保持终端精度?
  • RQ2该制导策略在无需重新训练的情况下,对实时传感器误差和执行器退化有多强的适应能力?
  • RQ3系统是否能有效在机动过程中偏转至新目标,同时保持终端精度和约束合规性?
  • RQ4通过延迟最终目标披露的多阶段偏转策略,武器在多大程度上可成功规避拦截器?
  • RQ5在完整的6-DOF动态环境中,制导与控制的集成性能如何?

主要发现

  • 基于强化学习的制导系统成功适应了非 nominal 条件,包括气动系数±20%的扰动和完全的执行器故障。
  • 系统始终符合所有关键飞行约束,包括热流率、动压、载荷因子和最小撞击速度。
  • 智能体在模拟中对机动地面目标实现了高精度终端打击,撞击误差始终低于5米。
  • 系统展示了有效的目标偏转能力,使武器能在保持终端性能和约束遵守的前提下转向新目标。
  • 采用多阶段偏转的威胁规避策略成功欺骗了具备有限中段修正能力的拦截器,显著提高了生存概率。
  • 初步的6-DOF仿真结果证实了制导与控制集成的稳定性和准确性,验证了端到端方法的可行性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。