Skip to main content
QUICK REVIEW

[论文解读] A Survey on Reinforcement Learning in Aviation Applications

Pouria Razzaghi, Amin Tabrizian|arXiv (Cornell University)|Nov 3, 2022
Air Traffic Management and Optimization被引用 22
一句话总结

对强化学习方法及其在航空领域应用的全面综述,涵盖标准RL公式、基于模型与无模型方法、演员-评论家与多智能体RL,以及选定的航空应用如碰撞避免、ATFM、ARM和飞行控制。

ABSTRACT

Compared with model-based control and optimization methods, reinforcement learning (RL) provides a data-driven, learning-based framework to formulate and solve sequential decision-making problems. The RL framework has become promising due to largely improved data availability and computing power in the aviation industry. Many aviation-based applications can be formulated or treated as sequential decision-making problems. Some of them are offline planning problems, while others need to be solved online and are safety-critical. In this survey paper, we first describe standard RL formulations and solutions. Then we survey the landscape of existing RL-based applications in aviation. Finally, we summarize the paper, identify the technical gaps, and suggest future directions of RL research in aviation.

研究动机与目标

  • 在航空情境下解释RL问题的形式化以及关键概念。
  • 综述无模型RL方法(基于价值、基于策略、演员-评论家及其扩展)及其延展。
  • 讨论多智能体RL及其与航空应用的相关性。
  • 回顾在航空领域如碰撞规避、空中交通流量管理、收益管理和飞行控制等方面的选定RL应用。
  • 指出差距并提出航空领域RL研究的未来方向。

提出的方法

  • 给出使用马尔可夫决策过程(MDP)的标准RL公式,并讨论基于模型与无模型方法。
  • 描述基于价值的方法(Q学习、DQN及其变体)及其深度扩展。
  • 解释基于策略的方法(REINFORCE、DPG、DDPG、PPO、TRPO)及演员-评论家混合方法。
  • 概述多智能体RL(MARL)框架及集中式/分布式训练范式。
  • 回顾航空领域RL应用的分类法,结合示例研究与方法。

实验结果

研究问题

  • RQ1哪些RL形式化与算法最适合航空领域的序贯决策问题?
  • RQ2基于模型、无模型以及演员-评论家方法在各航空领域的应用情况如何?
  • RQ3在航空情境中,MARL的角色与有效性如何?
  • RQ4哪些关键挑战(验证、仿真到现实的差距、样本效率、可解释性)阻碍RL在航空领域的现实应用?
  • RQ5哪些未来方向能够推动航空应用的RL研究?

主要发现

  • RL为航空领域的序贯决策问题提供了数据驱动的框架,包括安全关键的在线决策。
  • 在航空任务中已探索了广泛的RL方法——基于价值、基于策略、演员-评论家和MARL。
  • 碰撞避免、ATFM、ARM和飞行控制是活跃的RL应用领域,算法选择多样(DQN、PPO、DDPG、MADDPG、SAC)。
  • 挑战包括模型验证、从仿真到现实的差距、样本效率,以及对基于DRL的控制器的可解释性。
  • 该综述强调需要对RL进行正式验证,并为现实世界航空部署提供灵活、鲁棒的RL解决方案。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。