Skip to main content
QUICK REVIEW

[论文解读] A 3D Game Theoretical Framework for the Evaluation of Unmanned Aircraft Systems Airspace Integration Concepts

Negin Musavi, Ayman Manzoor|arXiv (Cornell University)|Feb 20, 2018
Aerospace and Aviation Technology参考文献 40被引用 3
一句话总结

本文提出了一种三维博弈论仿真框架,通过动态层级k推理与神经网络拟合Q学习(Neural Fitted Q-Learning)建模动态的人类飞行员决策行为,以评估无人机系统(UAS)在国家空域系统(NAS)中的融合。结果表明,有人驾驶飞机与UAS之间共享冲突解决责任可最小化间隔违反情况,并优化整体性能,优于仅由单一平台负责的场景。

ABSTRACT

Predicting the outcomes of integrating Unmanned Aerial Systems (UAS) into the National Airspace System (NAS) is a complex problem which is required to be addressed by simulation studies before allowing the routine access of UAS into the NAS. This paper focuses on providing a 3-dimensional (3D) simulation framework using a game theoretical methodology to evaluate integration concepts using scenarios where manned and unmanned air vehicles co-exist. In the proposed method, human pilot interactive decision making process is incorporated into airspace models which can fill the gap in the literature where the pilot behavior is generally assumed to be known a priori. The proposed human pilot behavior is modeled using dynamic level-k reasoning concept and approximate reinforcement learning. The level-k reasoning concept is a notion in game theory and is based on the assumption that humans have various levels of decision making. In the conventional "static" approach, each agent makes assumptions about his or her opponents and chooses his or her actions accordingly. On the other hand, in the dynamic level-k reasoning, agents can update their beliefs about their opponents and revise their level-k rule. In this study, Neural Fitted Q Iteration, which is an approximate reinforcement learning method, is used to model time-extended decisions of pilots with 3D maneuvers. An analysis of UAS integration is conducted using an example 3D scenario in the presence of manned aircraft and fully autonomous UAS equipped with sense and avoid algorithms.

研究动机与目标

  • 为解决UAS空域融合仿真中飞行员行为建模缺乏现实性的问题,现有研究常假设飞行员行为具有确定性。
  • 克服以往二维模型及混合空域仿真框架中静态策略假设的局限性。
  • 开发一个三维仿真环境,以捕捉不同认知水平飞行员的时间扩展型自适应决策行为。
  • 评估冲突解决责任分配(有人驾驶飞机 vs. UAS)对UAS-NAS融合中安全与性能指标的影响。
  • 通过行为感知的仿真框架,实现对UAS避撞算法、间隔标准与系统参数的量化评估。

提出的方法

  • 采用动态层级k推理建模异质飞行员认知水平(从层级0到层级2),使智能体在重复交互过程中可更新信念并调整策略。
  • 整合神经网络拟合Q学习(NFQ),在三维空域中建模时间扩展型、序列化决策行为,以深层函数逼近替代静态Q表。
  • 模拟混合空域场景,其中有人驾驶飞机采用动态层级k策略,而全自主UAS则配备SAA1与SAA2避撞算法。
  • 将冲突解决责任建模为可配置参数:仅有人驾驶飞机负责、仅UAS负责,或共享责任,各智能体遵循对应机动规则。
  • 采用基于最小间隔标准的三维空域模型与冲突检测机制。
  • 使用安全指标(间隔违反次数)与性能指标(航迹偏差与飞行时间)评估系统结果。

实验结果

研究问题

  • RQ1相较于静态或确定性模型,动态层级k推理在UAS-NAS融合仿真中如何提升飞行员行为建模的现实性与预测能力?
  • RQ2在三维单次相遇场景中,将冲突解决责任分配给不同平台(有人驾驶飞机、UAS或双方)对安全与性能有何影响?
  • RQ3在不同责任分配与飞行员认知水平分布下,SAA1与SAA2避撞算法的表现如何?
  • RQ4神经网络拟合Q学习在不依赖大型Q表的前提下,多大程度上实现了三维空域中自适应、时间扩展型决策的可扩展建模?
  • RQ5在共享冲突解决场景中,不同飞行员认知水平分布(10%层级0、60%层级1、30%层级2)对整体系统安全与性能有何影响?

主要发现

  • 在SAA1与SAA2两种算法下,共享冲突解决责任场景的间隔违反次数最少,有效最小化了安全风险。
  • 当仅由有人驾驶飞机负责冲突解决时,其平均航迹偏差高于共享责任场景,表明工作负荷增加且冲突解决效率降低。
  • 当UAS单独负责时,其平均航迹偏差高于共享责任场景,表明在单一责任下存在效率低下或机动策略次优的问题。
  • 当仅由有人驾驶飞机负责冲突解决时,UAS飞行时间最短,表明当UAS无需机动时可保持最优飞行路径。
  • 动态层级k推理模型成功捕捉了飞行员的自适应行为,层级1与层级2飞行员在重复交互中能更新信念并调整策略。
  • 神经网络拟合Q学习的集成实现了无需存储大型Q表的可扩展函数化策略学习,支持对复杂三维、时间扩展型决策序列的建模。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。