Skip to main content
QUICK REVIEW

[论文解读] Efficient Inference Without Trading-off Regret in Bandits: An Allocation Probability Test for Thompson Sampling

Nina Deliu, Joseph Jay Williams|arXiv (Cornell University)|Oct 29, 2021
Advanced Bandit Algorithms Research被引用 4
一句话总结

本文提出了一种名为分配概率检验(AP-test)的新颖假设检验方法,用于多臂赌博机数据,该方法利用Thompson Sampling的分配概率,实现在不牺牲遗憾性能的前提下进行有效推断。该方法在小样本中保持高统计效能,并控制第一类错误,且无需限制赌博机的探索行为或依赖大样本量。

ABSTRACT

Using bandit algorithms to conduct adaptive randomised experiments can minimise regret, but it poses major challenges for statistical inference (e.g., biased estimators, inflated type-I error and reduced power). Recent attempts to address these challenges typically impose restrictions on the exploitative nature of the bandit algorithm$-$trading off regret$-$and require large sample sizes to ensure asymptotic guarantees. However, large experiments generally follow a successful pilot study, which is tightly constrained in its size or duration. Increasing power in such small pilot experiments, without limiting the adaptive nature of the algorithm, can allow promising interventions to reach a larger experimental phase. In this work we introduce a novel hypothesis test, uniquely based on the allocation probabilities of the bandit algorithm, and without constraining its exploitative nature or requiring a minimum experimental size. We characterise our $Allocation\\ Probability\\ Test$ when applied to $Thompson\\ Sampling$, presenting its asymptotic theoretical properties, and illustrating its finite-sample performances compared to state-of-the-art approaches. We demonstrate the regret and inferential advantages of our approach, particularly in small samples, in both extensive simulations and in a real-world experiment on mental health aspects.

研究动机与目标

  • 为解决在通过最小遗憾赌博机算法(如Thompson Sampling)收集的数据上进行有效统计推断的挑战。
  • 克服现有方法的局限性:这些方法要么限制赌博机的探索行为(从而增加遗憾),要么需要大样本量以保证渐近有效性。
  • 开发一种假设检验方法,使其在小样本试点研究中保持高统计效能,同时确保第一类错误控制。
  • 为临床试验和移动健康等领域的应用研究人员提供一种实用的、非渐近的解决方案,这些领域常见小样本自适应实验。

提出的方法

  • AP-test基于实验臂的分配概率超过反映等概率随机化的阈值(π^ER = 1/(K+1))的时间比例。
  • 其检验统计量定义为:在时间步数T中,实验臂的估计分配概率超过π^ER的时间步数。
  • 通过蒙特卡洛模拟在每个时间步估计分配概率,以确保计算可行性。
  • 该检验使用原假设下检验统计量的精确分布来计算临界值并控制第一类错误,避免依赖渐近正态性。
  • 通过收敛性结果建立理论保证:在原假设下,检验统计量保持随机有界;在备择假设下,随着T增加,其发散至无穷大。
  • 该方法对基础奖励分布保持无偏性,且无需分批或分层的数据结构。

实验结果

研究问题

  • RQ1能否构建一种假设检验方法,使其在不限制赌博机探索行为的前提下,控制小样本赌博机实验中的第一类错误?
  • RQ2在小样本中,AP-test相较于BOLS和AW-AIPW等现有方法是否具有更高的统计效能?
  • RQ3在Thompson Sampling下,实验臂的分配概率如何演变?这一特性能否用于推断?
  • RQ4AP-test在标准Thompson Sampling下是否具有渐近有效性?当实验臂表现更优时,其是否收敛于拒绝原假设?

主要发现

  • 与AW-AIPW不同,AP-test在小样本中能有效控制第一类错误,而AW-AIPW在样本量较小时会夸大第一类错误。
  • AP-test即使在小批次规模下仍保持高统计效能——而BOLS在批次大小为3时效能低于0.1。
  • 在备择假设下,随着T增加,AP-test统计量发散至无穷大,证实了其的一致性。
  • 在时间推移中,Thompson Sampling下最优臂的分配概率以概率1收敛至1,这为检验的理论基础提供了支持。
  • 该方法可在不约束赌博机遗憾性能的前提下实现有效推断,从而允许对最优臂进行完全利用。
  • 模拟实验与一项真实世界心理健康实验的实证结果均表明,AP-test在小样本场景下优于当前最先进的替代方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。