Skip to main content
QUICK REVIEW

[论文解读] Learning to detect an oddball target with observations from an exponential family

Gayathri R Prabhu, Srikrishna Bhashyam|arXiv (Cornell University)|Dec 11, 2017
Advanced Bandit Algorithms Research参考文献 9被引用 9
一句话总结

本文提出一种序列策略,用于在多臂赌博机中检测一个参数未知的奇异常臂,该赌博机属于向量指数族分布,采用修正的广义似然比检验并结合共轭先验,以确保误检概率有界。该策略在最小化总成本(期望时间加上切换成本)方面渐近最优,其规模为 log(1/α)/D*,其中 α 为误检阈值,D* 为基于相对熵的最优缩放因子。

ABSTRACT

The problem of detecting an odd arm from a set of K arms of a multi-armed bandit, with fixed confidence, is studied in a sequential decision-making scenario. Each arm's signal follows a distribution from a vector exponential family. All arms have the same parameters except the odd arm. The actual parameters of the odd and non-odd arms are unknown to the decision maker. Further, the decision maker incurs a cost for switching from one arm to another. This is a sequential decision making problem where the decision maker gets only a limited view of the true state of nature at each stage, but can control his view by choosing the arm to observe at each stage. Of interest are policies that satisfy a given constraint on the probability of false detection. An information-theoretic lower bound on the total cost (expected time for a reliable decision plus total switching cost) is first identified, and a variation on a sequential policy based on the generalised likelihood ratio statistic is then studied. Thanks to the vector exponential family assumption, the signal processing in this policy at each stage turns out to be very simple, in that the associated conjugate prior enables easy updates of the posterior distribution of the model parameters. The policy, with a suitable threshold, is shown to satisfy the given constraint on the probability of false detection. Further, the proposed policy is asymptotically optimal in terms of the total cost among all policies that satisfy the constraint on the probability of false detection.

研究动机与目标

  • 解决在多臂赌博机中检测奇异常臂的问题,其中各臂的真实参数未知,且观测值来自向量指数族分布。
  • 将切换成本纳入序列检测框架,其中每次切换臂均产生成本。
  • 设计一种策略,在保证误检概率的置信度约束下,最小化总成本(期望时间加上切换成本)。
  • 将先前针对泊松分布观测值的研究推广至更广泛的指数族分布类别。
  • 建立总成本的信息论下界,并证明所提策略渐近达到该下界。

提出的方法

  • 采用修正的广义似然比检验(GLRT),其中似然比在未知参数的虚拟共轭先验上进行平均,以稳定检验统计量。
  • 对指数族使用共轭先验,以实现在每个阶段的高效且闭式表达的后验更新。
  • 基于对数似然比统计量采用时不变阈值策略,以决定何时停止采样并宣告奇异常臂。
  • 利用相对熵和对数归一化函数的共轭凸函数,推导出总成本的信息论下界。
  • 设计一种采样策略,通过根据相对熵所导出的最优比例分配观测值,使总成本渐近匹配下界。
  • 基于广义似然比统计量设计阈值规则,以确保误检概率被控制在 α 以内。

实验结果

研究问题

  • RQ1在固定置信度约束下,检测奇异常臂的总成本(期望时间加切换成本)是否存在基本下界?
  • RQ2当指数族分布的参数未知时,如何设计一种序列策略,使其渐近达到该下界?
  • RQ3共轭先验在实现所提策略中高效且可计算的后验更新中起到何种作用?
  • RQ4引入切换成本如何影响最优采样策略和总成本?
  • RQ5能否证明所提策略在所有满足误检约束的策略中,于总成本方面渐近最优?

主要发现

  • 所提策略渐近达到总成本的信息论下界,总成本的规模为 log(1/α)/D*,其中 D* 为最优相对熵基缩放因子。
  • 通过在广义似然比统计量上精心选择阈值,策略确保了误检概率被控制在 α 以内。
  • 该策略的采样策略在目标误检概率 α 趋近于零时,收敛至下界所建议的最优分配比例。
  • 共轭先验的使用使得后验更新简单高效,使策略在实时实现中具有计算可操作性。
  • 该策略在所有满足误检约束的策略中,于最小化期望时间与总切换成本之和方面渐近最优。
  • 该分析将先前针对泊松分布观测值的结果推广至完整的向量指数族分布类别,统一了高斯、二项分布和伽马分布等模型。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。