[论文解读] Optimality of Thompson Sampling for Gaussian Bandits Depends on Priors
该论文表明,对于均值和方差未知的高斯老虎机问题,Thompson Sampling (TS) 仅在特定先验下才能实现渐近最优性——具体而言,仅当使用均匀先验时,才能达到理论 regret 上界;而 Jeffreys 先验和参考先验则无法实现,导致多项式 regret。研究表明,在多参数模型中,先验选择对 TS 性能具有决定性影响。
In stochastic bandit problems, a Bayesian policy called Thompson sampling (TS) has recently attracted much attention for its excellent empirical performance. However, the theoretical analysis of this policy is difficult and its asymptotic optimality is only proved for one-parameter models. In this paper we discuss the optimality of TS for the model of normal distributions with unknown means and variances as one of the most fundamental example of multiparameter models. First we prove that the expected regret of TS with the uniform prior achieves the theoretical bound, which is the first result to show that the asymptotic bound is achievable for the normal distribution model. Next we prove that TS with Jeffreys prior and reference prior cannot achieve the theoretical bound. Therefore the choice of priors is important for TS and non-informative priors are sometimes risky in cases of multiparameter models.
研究动机与目标
- 研究在均值和方差均未知的正态分布下,多臂老虎机问题中 Thompson Sampling (TS) 的渐近最优性。
- 确定 TS 是否能实现该基本多参数模型的理论期望 regret 下界。
- 评估不同非信息先验(均匀先验、Jeffreys 先验、参考先验)对 TS regret 性能的影响。
- 澄清在多参数设置中,非信息先验是否安全或存在风险,尤其是在期望实现渐近最优性时。
提出的方法
- 形式化定义在均值和方差未知的正态分布下的 K-臂随机老虎机问题。
- 将期望 regret 定义为随时间累积的最优奖励与实际奖励之间的差值。
- 使用共轭先验进行贝叶斯更新,以建模对均值和方差参数的后验分布。
- 应用大偏差理论和 Chernoff 不等式,分析样本均值和平方和的尾部概率。
- 基于先验分布,推导出子优臂被选中的后验概率的上界。
- 证明:均匀先验可实现与理论下界匹配的对数 regret,而 Jeffreys 先验和参考先验则导致多项式 regret。
实验结果
研究问题
- RQ1Thompson Sampling 是否能为均值和方差未知的高斯老虎机问题实现理论上的渐近 regret 上界?
- RQ2在该多参数模型中,先验选择(均匀先验、Jeffreys 先验或参考先验)如何影响 Thompson Sampling 的 regret 性能?
- RQ3尽管应用广泛,非信息先验(如 Jeffreys 先验或参考先验)是否可能在多参数老虎机问题中导致次优 regret?
- RQ4即使先验是非信息性的,TS 的渐近最优性是否仍对先验设定敏感?
- RQ5在不同先验下,期望 regret 的精确行为如何,特别是其增长速率(对数型 vs. 多项式型)?
主要发现
- 使用均匀先验的 Thompson Sampling 在均值和方差未知的高斯老虎机问题中,可实现理论上的渐近 regret 上界。
- 使用 Jeffreys 先验的 Thompson Sampling 无法达到理论上的上界,反而在期望下表现出多项式 regret。
- 使用参考先验的 Thompson Sampling 同样无法达到理论上的上界,表现出类似的次优 regret 行为。
- 在多参数模型中,先验的选择显著影响 Thompson Sampling 的性能,即使这些先验是非信息性的。
- 本研究揭示,在多参数设置中,非信息先验可能具有风险,因为它们可能导致后验信念过于乐观,从而造成探索不足。
- 理论分析确认,只有特定先验(如均匀先验)才能确保此类老虎机问题中的渐近最优性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。