Skip to main content
QUICK REVIEW

[论文解读] A Pursuit-Evasion Differential Game with Strategic Information Acquisition

Yunhan Huang, Quanyan Zhu|arXiv (Cornell University)|Feb 10, 2021
Guidance and Control Systems参考文献 38被引用 7
一句话总结

本文提出了追逃暴露隐蔽(PEEC)博弈,一种线性二次高斯微分博弈,其中参与者战略性地决定何时相互观测,需承担直接观测成本以及状态暴露的间接成本。作者使用变分法和配方平方法分别推导出纳什控制与观测策略,表明机动性较低的参与者更倾向于隐蔽,且在无限时域极限下周期性观测成为最优策略,使期望跟踪误差趋近于零且二阶矩有界。

ABSTRACT

This paper studies a two-person linear-quadratic-Gaussian pursuit-evasion differential game with costly but controlled information. One player can decide when to observe the other player's state. However, one observation of another player's state comes with two costs: the direct cost of observing and the implicit cost of exposing his state. We call games of this type a Pursuit-Evasion-Exposure-Concealment (PEEC) game. The PEEC game constitutes two types of strategies: The control strategies and the observation strategies. We fully characterize the Nash control strategies of the PEEC game using techniques such as completing squares and the calculus of variations. We show that the derivation of the Nash observation strategies and the Nash control strategies can be decoupled. We develop a set of necessary conditions that facilitate the numerical computation of the Nash observation strategies. We show, in theory, that players with less maneuverability prefer concealment to exposure. We also show that when the game's horizon goes to infinity, the Nash observation strategy is to observe periodically, and the expected distance between the pursuer and the evader goes to zero with a bounded second moment. We conducted a series of numerical experiments to study the proposed PEEC game. We illustrate the numerical results using both figures and animation. Numerical results show that the pursuer can maintain high-grade performance even when the number of observations is limited. We also show that an evader with low maneuverability can still escape if the evader increases his stealthiness.

研究动机与目标

  • 建模观测成本高昂且暴露观测者状态的追逃博弈,捕捉现实世界中传感与隐身之间的权衡。
  • 构建一个博弈论框架——称为PEEC——在信息成本不对称下整合控制与观测决策。
  • 表征控制与观测的纳什均衡策略,通过解耦推导以提升可计算性。
  • 分析机动性与观测成本对参与者最优行为的影响,特别是隐蔽与暴露之间的权衡。
  • 建立最优观测时机的理论条件,并证明在无限时域情况下收敛至周期性策略。

提出的方法

  • 构建一个有限时域的双人线性二次高斯微分博弈,纳入观测成本与状态暴露惩罚。
  • 使用配方平方法与变分法推导出与观测决策无关的显式纳什控制策略。
  • 将观测策略的纳什推导与控制策略解耦,实现独立优化。
  • 利用一阶最优性条件与莱布尼茨法则对代价泛函进行分析,推导最优观测时机的必要条件。
  • 提出基于必要条件的数值计算框架,以计算最优观测时刻。
  • 分析无限时域极限,证明最优观测策略趋于周期性,且追逃者间期望距离收敛至零,二阶矩有界。

实验结果

研究问题

  • RQ1在追逃设定中,参与者如何最优平衡观测对手的成本与自身状态暴露的风险?
  • RQ2纳什控制与观测策略能否独立推导?这种解耦在何种条件下成立?
  • RQ3参与者的机动性对其偏好隐蔽还是暴露有何影响?
  • RQ4随着博弈时域趋于无限,最优观测策略是否收敛至周期性采样?
  • RQ5在最优策略下,期望跟踪误差随时间如何演变?其长期有界性如何?

主要发现

  • 无论追逃者的策略如何,逃逸者的最优观测策略始终为不观测,原因在于观测成本与暴露风险。
  • 追逃者的最优策略最小化一个包含估计误差迹与观测成本的代价泛函,最优观测时刻满足一阶必要条件。
  • 机动性较低的参与者更偏好隐蔽而非暴露,因其被探测后更难逃脱。
  • 在无限时域极限下,最优观测策略变为周期性,追逃者间期望距离收敛至零,且二阶矩有界。
  • 数值实验表明,即使观测次数有限,追逃者仍能保持高性能;而机动性较低的逃逸者可通过增强隐蔽性实现逃脱。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。