Skip to main content
QUICK REVIEW

[论文解读] Optimal Solutions for Adaptive Search Problems with Entropy Objectives

Hanyi Ding, David A. Castañón|arXiv (Cornell University)|Aug 17, 2015
Advanced Bandit Algorithms Research参考文献 32被引用 3
一句话总结

该论文将基于熵的自适应搜索问题建模为随机控制问题,通过凸优化推导出最优传感器管理策略。针对单传感器和多传感器场景,提供了可构造的算法,表明最优策略可通过标量凸规划计算得出,并在对称条件下实现解析解,同时对有限时域搜索提供了精确的基于熵的成本表征。

ABSTRACT

The problem of searching for an unknown object occurs in important applications ranging from security, medicine and defense. Sensors with the capability to process information rapidly require adaptive algorithms to control their search in response to noisy observations. In this paper, we discuss classes of dynamic, adaptive search problems, and formulate the resulting sensor control problems as stochastic control problems with imperfect information, based on previous work on noisy search problems. The structure of these problems, with objective functions related to information entropy, allows for a complete characterization of the optimal strategies and the optimal cost for the resulting finite-horizon stochastic control problems. We study the problem where an individual sensor is capable of searching over multiple sub-regions in a time, and provide a constructive algorithm for determining optimal policies in real time based on convex optimization. We also study the problem in which there are multiple sensors, each of which is only capable of detecting over one sub-region in a time, jointly searching for an object. Whereas this can be viewed as a special case of our multi-region results, we show that the computation of optimal policies can be decoupled into single-sensor individual scalar convex optimization problems, and provide simple symmetry conditions where the solutions can be determined analytically. We also consider the case where individual sensors can select the accuracy of their sensing modes with different costs, and derive optimal strategies for these problems in terms of the solutions of scalar convex optimization problems. We illustrate our results with experiments using multiple sensors searching for a single object.

研究动机与目标

  • 为具有噪声二值观测的传感器在动态、有限时域设置下,开发最优自适应搜索策略。
  • 将传感器控制建模为基于后验熵的信息论目标的随机控制问题。
  • 通过凸优化技术实现实时策略计算,适用于单传感器和多传感器配置。
  • 在对称条件下,将多传感器搜索解耦为独立的凸优化问题,实现解析解。
  • 将感知模式选择(含可变精度与成本)纳入模型,通过凸规划推导最优权衡。

提出的方法

  • 将自适应搜索问题建模为具有不完美信息和基于熵的成本函数的有限时域随机控制问题。
  • 使用后验微分熵作为目标,以最小化随时间推移的对象定位不确定性。
  • 通过动态规划推导最优策略,表明值函数依赖于初始熵和一个常数最优增益 $ G^* $。
  • 应用凸优化计算最优感知模式和工作点,利用香农熵的严格凹性。
  • 对于多传感器系统,当满足对称条件时,将联合优化问题分解为独立的标量凸规划。
  • 引入一个成本加权的熵增益函数 $ G(oldsymbol{u}, oldsymbol{l}) $,用于平衡信息增益与感知成本 $ eta $,实现权衡优化。

实验结果

研究问题

  • RQ1对于具有噪声二值观测的单传感器,如何获得最小化后验熵的最优自适应搜索策略?
  • RQ2在基于熵的目标下,如何实现实时计算有限时域搜索的最优传感器管理策略?
  • RQ3在对称条件下,多传感器搜索问题能否被分解为独立的优化问题?
  • RQ4在自适应搜索中,感知精度、成本与信息增益之间的最优权衡是什么?
  • RQ5最优成本结构如何与熵减少及感知模式选择相关联?

主要发现

  • 有限时域搜索的最优成本为 $ V(p_n, n) = H(p_n) - (N - n)G^* $,其中 $ G^* $ 为每步的最大信息增益。
  • 当联合感知模式 $ oldsymbol{A}_{n+1} $ 和工作点 $ oldsymbol{l}_{n+1} $ 达到最优 $ (oldsymbol{u}^*, oldsymbol{l}^*) $ 时,实现最优策略,此时 $ G(oldsymbol{u}, oldsymbol{l}) $ 取得最大值。
  • 对于对称多传感器系统,$ G^* = igoplus_{m=1}^M G^{(m)*} $,使得可通过单个传感器优化实现最优策略的解耦计算。
  • 最优联合感知模式可表示为边缘分布的乘积 $ ar{u}_{i_{1:M}} = igotimes_{m=1}^M (u^{(m)*})^{i_m}(1 - u^{(m)*})^{1-i_m} $,从而实现 $ G^* = igoplus_{m=1}^M G^{(m)*} $。
  • 当传感器可选择不同精度的感知模式并伴随相应成本时,最优策略由成本加权增益 $ G(oldsymbol{u}, oldsymbol{l}) $ 的标量凸优化导出。
  • 本文证明了最优策略存在且可通过凸规划实现实时计算,避免了难以处理的动态规划求解。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。