[论文解读] Optimal Event-Driven Multi-Agent Persistent Monitoring of a Finite Set of Targets
该论文提出了一种基于事件触发的、梯度驱动的控制框架,采用无穷小扰动分析(IPA)来优化一维空间中有限个目标的多智能体轨迹,以实现持续监控。通过将最优控制问题重新表述为关于切换点和驻留时间的参数优化,并引入基于势场的代价度量以确保梯度激励,该方法实现了鲁棒的在线近似最优不确定性减少,仿真结果验证了其性能接近基于图的调度方法。
We consider the problem of controlling the movement of multiple cooperating agents so as to minimize an uncertainty metric associated with a finite number of targets. In a one-dimensional mission space, we adopt an optimal control framework and show that the solution is reduced to a simpler parametric optimization problem: determining a sequence of locations where each agent may dwell for a finite amount of time and then switch direction. This amounts to a hybrid system which we analyze using Infinitesimal Perturbation Analysis (IPA) to obtain a complete on-line solution through an event-driven gradient-based algorithm which is also robust with respect to the uncertainty model used. The resulting controller depends on observing the events required to excite the gradient-based algorithm, which cannot be guaranteed. We solve this problem by proposing a new metric for the objective function which creates a potential field guaranteeing that gradient values are non-zero. This approach is compared to an alternative graph-based task scheduling algorithm for determining an optimal sequence of target visits. Simulation examples are included to demonstrate the proposed methods.
研究动机与目标
- 解决在一条一维任务空间中,使用多个协作智能体对有限数量动态目标进行持续监控的问题。
- 通过最优智能体轨迹控制,最小化与目标状态相关的系统级不确定性度量。
- 克服标准IPA公式中因事件激励不足而导致的梯度优化失败问题。
- 开发一种可扩展的在线控制策略,在随机不确定性模型下仍保持鲁棒性。
- 将所提出的事件驱动IPA方法与基于图的任务调度方法进行性能比较。
提出的方法
- 将持续监控问题建模为基于目标不确定性的代价函数的最优控制问题。
- 将最优控制问题简化为关于切换点和驻留时间的参数优化,将智能体轨迹表征为混合系统。
- 应用无穷小扰动分析(IPA)在线计算代价函数相对于轨迹参数的梯度。
- 引入改进的代价度量 $ J_2(\bm{\theta}, \bm{\omega}, t) $,构建势场,确保即使智能体初始时刻错过目标,也能保持非零梯度和事件激励。
- 使用由事件触发的IPA驱动的梯度下降算法,迭代调整轨迹以逼近最优解。
- 在确定性和随机环境下通过仿真验证该方法,包括随机目标动态和位置。
实验结果
研究问题
- RQ1能否将一维空间中多智能体持续监控的最优控制问题,简化为关于切换点和驻留时间的参数优化?
- RQ2当目标访问事件未发生导致梯度为零时,如何使基于IPA的梯度估计保持鲁棒性?
- RQ3与基于图的任务调度方法相比,所提出的事件驱动IPA方法在性能上有哪些提升?
- RQ4所提出的势场代价度量如何在未发生目标访问的情况下确保梯度激励和收敛性?
- RQ5在目标动态和位置存在随机不确定性模型时,基于IPA的控制器在多大程度上保持鲁棒性?
主要发现
- 最优智能体轨迹被限制在目标位置的凸包内,证实了理论收敛边界。
- 在双智能体、五目标的仿真中,基于IPA的控制器最终代价为4.99,与基于图的方法的4.92非常接近。
- 在不确定性流入速率服从均匀分布($A_i \sim U(0,2)$)的随机仿真中,IPA代价为42.46,接近确定性最优代价29.40。
- 当目标位置随机扰动($\sim U(x_i \pm 0.25)$)时,IPA代价为34.89,同样接近确定性最优值。
- 引入$J_2$势场代价度量后,即使初始轨迹远离目标,系统也能实现收敛,代价从30.24在100次迭代后降至30.24。
- 该方法对随机不确定性表现出鲁棒性,收敛速度取决于随机过程的方差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。