Skip to main content
QUICK REVIEW

[论文解读] Asymptotically Optimal Sequential Experimentation Under Generalized Ranking

Wesley Cowan, Michael N. Katehakis|arXiv (Cornell University)|Oct 7, 2015
Advanced Bandit Algorithms Research参考文献 30被引用 4
一句话总结

本文提出了一种广义框架,用于多臂赌博机问题中的序列实验,支持任意评分函数,超越基于均值的优化。它建立了次优采样数的渐近下界,并证明在温和正则性条件下,采用特定置信区间的一类UCB策略可达到该下界,从而确保在帕累托、均匀和正态分布下对多种评分函数的渐近最优性。

ABSTRACT

We consider the \mnk{classical} problem of a controller activating (or sampling) sequentially from a finite number of $N \geq 2$ populations, specified by unknown distributions. Over some time horizon, at each time $n = 1, 2, \ldots$, the controller wishes to select a population to sample, with the goal of sampling from a population that optimizes some "score" function of its distribution, e.g., maximizing the expected sum of outcomes or minimizing variability. We define a class of extit{Uniformly Fast (UF)} sampling policies and show, under mild regularity conditions, that there is an asymptotic lower bound for the expected total number of sub-optimal population activations. Then, we provide sufficient conditions under which a UCB policy is UF and asymptotically optimal, since it attains this lower bound. Explicit solutions are provided for a number of examples of interest, including general score functionals on unconstrained Pareto distributions (of potentially infinite mean), and uniform distributions of unknown support. Additional results on bandits of Normal distributions are also provided.

研究动机与目标

  • 解决在未知分布的 N 个总体中进行序列采样时,当最优选择由广义评分函数而非仅期望值定义的挑战。
  • 为任意一致快速(UF)策略下次优激活次数的期望值建立理论下界。
  • 确定UCB型策略达到该下界的充分条件,从而证明渐近最优性。
  • 将赌博机算法的适用范围扩展至重尾或期望为无穷的分布,如帕累托分布和未知支持区间的均匀分布。
  • 为实际应用提供显式解法和有限时域边界,适用于多种分布族和评分函数。

提出的方法

  • 引入一类一致快速(UF)采样策略,确保次优激活次数随时间呈次多项式增长。
  • 定义一种UCB-$(\mathcal{F},s,\tilde{d})$ 策略,其置信区间基于Kullback-Leibler散度和评分函数的实证估计。
  • 基于最优与次优分布之间的差异,建立次优激活期望数的一般渐近下界。
  • 使用Kullback-Leibler信息散度 $\mathbf{I}(f,g)$ 作为度量,量化分布间的统计分离程度,并推导置信区间。
  • 应用 R-条件(R1′, R2′, R3′)验证,在评分函数和分布族的温和正则性条件下,UCB策略可达到下界。
  • 推导出在特定模型下(包括正态、帕累托和均匀分布)UCB策略的次优拉动次数的显式渐近表达式。

实验结果

研究问题

  • RQ1当最优老虎机由广义评分函数而非均值定义时,序列实验中次优激活次数的根本极限是什么?
  • RQ2在何种条件下,基于UCB的策略可达到该根本下界,从而确保渐近最优性?
  • RQ3当评分函数不是均值(如无穷期望的帕累托分布)时,次优采样行为的渐近特性如何变化?
  • RQ4UCB框架能否扩展以处理未知支持的分布(如任意区间上的均匀分布)?
  • RQ5当评分函数为尾概率 $\mathbb{P}(X \geq \kappa)$ 时,正态赌博机中UCB策略的次优拉动的精确渐近速率是多少?

主要发现

  • 对于任意UF策略,次优激活次数的期望值增长慢于任何正幂次的 $n$,为性能设定了基准。
  • 基于最优与次优分布之间的Kullback-Leibler散度,推导出次优激活期望数的渐近下界。
  • 对于评分函数为 $s_{\kappa}(f) = \mathbb{P}(X \geq \kappa)$ 的正态赌博机模型,UCB策略 $\pi^{*}_{\kappa}$ 渐近最优,达到所推导的下界。
  • 对于基于阈值的评分函数的正态赌博机模型,次优拉动的渐近速率为 $\lim_{n \to \infty} \frac{\mathbb{E}[T^{i}_{\pi^{*}_{\kappa}}(n)]}{\ln n} = \frac{2}{\left(\frac{\kappa - \mu_i}{\sigma_i} - \Phi^{-1}(1 - s^*) \right)^2}$。
  • 在R-条件下,策略被证明渐近最优,通过KL散度和正态近似验证了连续性、集中性和尾部概率边界。
  • 该框架适用于帕累托、均匀和正态分布,为实际应用提供了显式渐近表达式和有限时域边界。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。