[论文解读] Integrated and Adaptive Guidance and Control for Endoatmospheric Missiles via Reinforcement Learning
本文提出了一种用于临近大气层导弹的强化元学习框架,实现集成化、自适应的制导与控制,将导航输出直接映射到控制面偏转。该系统在大飞行包线变化、非额定气动参数、柔性机体动力学以及导引头引起的姿态误差条件下均表现出鲁棒的拦截性能,优于采用三环自动驾驶仪的传统比例导航方法,在仿真中表现更优。
We apply a reinforcement meta-learning framework to optimize an integrated and adaptive guidance and flight control system for an air-to-air missile. The system is implemented as a policy that maps navigation system outputs directly to commanded rates of change for the missile's control surface deflections. The system induces intercept trajectories against a maneuvering target that satisfy control constraints on fin deflection angles, and path constraints on look angle and load. We test the optimized system in a six degrees-of-freedom simulator that includes a non-linear radome model and a strapdown seeker model, and demonstrate that the system adapts to both a large flight envelope and off-nominal flight conditions including perturbation of aerodynamic coefficient parameters and center of pressure locations, and flexible body dynamics. Moreover, we find that the system is robust to the parasitic attitude loop induced by radome refraction and imperfect seeker stabilization. We compare our system's performance to a longitudinal model of proportional navigation coupled with a three loop autopilot, and find that our system outperforms this benchmark by a large margin. Additional experiments investigate the impact of removing the recurrent layer from the policy and value function networks, performance with an infrared seeker, and flexible body dynamics.
研究动机与目标
- 开发一种适用于空空导弹的集成制导与控制系统,可适应大飞行包线变化和非额定条件。
- 解决舵面偏转的控制约束、视线角和载荷的路径约束,以及导引头引起的寄生姿态回路问题。
- 用一种端到端学习的策略替代传统制导律和自动驾驶仪,该策略将导航输出直接映射为控制面指令。
- 评估系统对气动系数扰动、压力中心偏移以及柔性体动力学的鲁棒性。
- 在高保真六自由度仿真环境中,证明其优于基准比例导航与三环自动驾驶仪的性能。
提出的方法
- 采用元学习框架训练深度强化学习策略,以实现对多样化飞行条件的泛化能力。
- 该策略将实时导航系统输出直接映射为导弹控制面偏转的指令变化率。
- 策略网络和价值函数网络中均引入循环神经网络,以建模控制动作中的时序依赖性。
- 六自由度仿真环境包含非线性整流罩效应和捷联导引头模型,以模拟真实的导引头动力学。
- 训练环境引入了气动系数、压力中心位置以及柔性体动力学的扰动,以测试系统的鲁棒性。
- 性能评估基于一个基准纵向比例导航系统与三环自动驾驶仪的对比。
实验结果
研究问题
- RQ1单一强化学习策略是否能在不同气动和动力学条件下,有效应对临近大气层导弹的全飞行包线?
- RQ2在策略网络和价值函数网络中引入循环层,对系统性能和适应性有何影响?
- RQ3系统在多大程度上能抵御由整流罩折射和导引头稳定不完善引起的寄生姿态回路?
- RQ4系统在柔性体动力学和非额定气动参数变化下的表现如何?
- RQ5所学习的策略在拦截成功率和约束满足方面是否优于经典的比例导航与三环自动驾驶仪?
主要发现
- 基于强化学习的系统成功生成了满足舵面偏转控制约束以及视线角和载荷路径约束的拦截轨迹。
- 系统在大飞行包线变化下表现出鲁棒性能,包括气动系数和压力中心位置的扰动。
- 尽管受到整流罩折射和导引头稳定不完善引起的寄生姿态回路影响,系统仍保持稳定和高效。
- 在策略和价值函数网络中引入循环层显著提升了性能,尤其在处理时序动态和适应变化条件方面。
- 系统在拦截成功率和约束遵守方面优于基准比例导航与三环自动驾驶仪。
- 针对红外导引头和柔性体动力学的实验验证了系统在不同传感器和机体构型下的适应性和鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。