Skip to main content
QUICK REVIEW

[论文解读] Quickest Search Over Multiple Sequences with Mixed Observations

Jun Geng, Weiyu Xu|arXiv (Cornell University)|Feb 15, 2013
Advanced Statistical Process Monitoring参考文献 17被引用 4
一句话总结

该论文提出了一种针对多个序列的两阶段最快搜索策略,其中观测值可以是来自不同序列样本的线性组合,从而实现对由分布 $F_1$ 生成的稀有序列的更快识别。通过在扫描阶段使用混合观测值高效剔除非目标序列,并在第二阶段进行精细化选择,该方法显著降低了搜索延迟——尤其当 $F_1$ 稀有时——与单观测方法相比,检测速度最高可提升 40%。

ABSTRACT

The problem of sequentially finding an independent and identically distributed (i.i.d.) sequence that is drawn from a probability distribution $F_1$ by searching over multiple sequences, some of which are drawn from $F_1$ and the others of which are drawn from a different distribution $F_0$, is considered. The sensor is allowed to take one observation at a time. It has been shown in a recent work that if each observation comes from one sequence, Cumulative Sum (CUSUM) test is optimal. In this paper, we propose a new approach in which each observation can be a linear combination of samples from multiple sequences. The test has two stages. In the first stage, namely scanning stage, one takes a linear combination of a pair of sequences with the hope of scanning through sequences that are unlikely to be generated from $F_1$ and quickly identifying a pair of sequences such that at least one of them is highly likely to be generated by $F_1$. In the second stage, namely refinement stage, one examines the pair identified from the first stage more closely and picks one sequence to be the final sequence. The problem under this setup belongs to a class of multiple stopping time problems. In particular, it is an ordered two concatenated Markov stopping time problem. We obtain the optimal solution using the tools from the multiple stopping time theory. Numerical simulation results show that this search strategy can significantly reduce the searching time, especially when $F_{1}$ is rare.

研究动机与目标

  • 解决在多个序列中快速识别由稀有分布 $F_1$ 生成的序列的问题,其中部分序列来自 $F_0$。
  • 通过允许观测值为多个序列样本的线性组合,而非仅单序列观测,提升搜索效率。
  • 在受限误差概率下,最小化搜索延迟与误差概率的线性组合。
  • 利用多停止时间理论,为两阶段序列搜索过程设计最优策略。

提出的方法

  • 该方法采用两阶段流程:扫描阶段使用混合观测,精炼阶段使用单序列观测。
  • 在扫描阶段,传感器对两个序列取线性组合,以检验两者是否均可能来自 $F_0$;若是,则将两者剔除。
  • 当传感器基于后验概率获得足够把握,确认至少一个序列来自 $F_1$ 时,扫描阶段停止。
  • 在精炼阶段,对候选对使用类似 CUSUM 的检验,选择最可能来自 $F_1$ 的序列,以最小化误差概率。
  • 利用有序两段连接的马尔可夫停止时间理论,推导出最优停止时间与切换规则。
  • 最优决策规则在精炼阶段选择后验概率更高、更可能由 $F_1$ 生成的序列。

实验结果

研究问题

  • RQ1使用多个序列的观测值线性组合,能否提升对稀有 $F_1$ 序列的最快搜索效率?
  • RQ2如何设计扫描与精炼阶段,以在搜索延迟与误差概率之间实现最优权衡?
  • RQ3当允许使用混合观测时,扫描阶段的最优停止与切换规则是什么?
  • RQ4与单观测策略相比,使用混合观测在搜索延迟与误差性能方面表现如何?
  • RQ5在具有马尔可夫状态动态的两阶段序列搜索中,最优策略的结构是怎样的?

主要发现

  • 当 $F_1$ 稀有时,所提出的混合观测策略相比单观测策略,将搜索延迟降低了约 40%。
  • 扫描阶段的最优停止规则为基于区域的规则,依赖于后验概率 $p^{1,1}$ 与 $p^{mix}$,当对至少一个 $F_1$ 序列的置信度足够高时停止。
  • 扫描阶段的最优切换规则同样为基于区域的规则,当两个序列均可能来自 $F_0$ 时触发序列替换。
  • 精炼阶段的最优停止时间出现在误差期望成本低于未来成本时,确保延迟与误差最小化。
  • 代价函数 $V_s(p^{1,1}, p^{mix})$ 为凹函数且有上界,平坦区域对应切换区域 $R_\theta$,证实了所推导策略的最优性。
  • 数值结果表明,该策略在先验概率较低($\pi = 0.05$)且信噪比(SNR)较高时表现最佳,且随着 SNR 增加,性能持续提升。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。