[论文解读] First-Order Bayesian Regret Analysis of Thompson Sampling
本文针对组合半-bandit问题中的Thompson采样提出了改进的信息论分析,引入了对尺度敏感的信息比和坐标熵,实现了半-bandit设置下首阶贝叶斯遗憾界$\widetilde{O}(\sqrt{dL^*})$,与目前已知的最佳频率学遗憾界一致。此外,本文提出了阈值化Thompson采样(Thresholded Thompson Sampling),在$L^* \leq \overline{L}^*$时实现了与$T$无关的遗憾界,这是标准Thompson采样所不具备的性质。
We address online combinatorial optimization when the player has a prior over the adversary's sequence of losses. In this framework, Russo and Van Roy proposed an information-theoretic analysis of Thompson Sampling based on the information ratio, resulting in optimal worst-case regret bounds. In this paper we introduce three novel ideas to this line of work. First we propose a new quantity, the scale-sensitive information ratio, which allows us to obtain more refined first-order regret bounds (i.e., bounds of the form $\sqrt{L^*}$ where $L^*$ is the loss of the best combinatorial action). Second we replace the entropy over combinatorial actions by a coordinate entropy, which allows us to obtain the first optimal worst-case bound for Thompson Sampling in the combinatorial setting. Finally, we introduce a novel link between Bayesian agents and frequentist confidence intervals. Combining these ideas we show that the classical multi-armed bandit first-order regret bound $ ilde{O}(\sqrt{d L^*})$ still holds true in the more challenging and more general semi-bandit scenario. This latter result improves the previous state of the art bound $ ilde{O}(\sqrt{(d+m^3)L^*})$ by Lykouris, Sridharan and Tardos. Moreover we sharpen these results with two technical ingredients. The first leverages a recent insight of Zimmert and Lattimore to replace Shannon entropy with more refined potential functions in the analysis. The second is a \emph{Thresholded} Thompson sampling algorithm, which slightly modifies the original algorithm by never playing low-probability actions. This thresholding results in fully $T$-independent regret bounds when $L^*$ is almost surely upper-bounded, which we show does not hold for ordinary Thompson sampling.
研究动机与目标
- 弥合组合半-bandit问题中贝叶斯与频率学遗憾界之间的差距。
- 发展一种精细化的信息论分析方法,用于Thompson采样,以捕捉与$L^*$(最优损失)相关的首阶遗憾缩放。
- 在最优损失有界$ L^* \leq \overline{L}^*$ 的条件下,建立与$T$无关的遗憾界,而标准Thompson采样无法实现此性质。
- 通过将贝叶斯代理与频率学置信区间关联,统一贝叶斯与频率学视角。
提出的方法
- 引入对尺度敏感的信息比,作为对信息比的改进,以捕捉与损失相关的遗憾缩放。
- 用坐标熵替代组合动作上的全局香农熵,以提升组合设置下的分析精度。
- 利用Tsallis熵和对数障碍势函数,将信息论分析从香农熵推广至更精细的框架。
- 提出阈值化Thompson采样,避免选择低概率动作,从而在有界$L^*$下保持鲁棒性。
- 采用新颖的贝叶斯假设检验论证,证明良好臂的后验概率存在统一下界。
- 应用Sion极小化极大定理,将问题转化为贝叶斯设置,从而在遗憾分析中使用先验分布。
实验结果
研究问题
- RQ1Thompson采样是否能在半-bandit组合设置中实现形式为$O(\sqrt{dL^*})$的首阶遗憾界?
- RQ2标准Thompson采样算法在满足$L^* \leq \overline{L}^*$几乎必然成立时,是否能实现与$T$无关的遗憾界?
- RQ3能否以有意义的方式将Thompson采样的贝叶斯分析与频率学置信区间联系起来?
- RQ4能否通过使用对尺度敏感或坐标特定的度量,对信息比框架进行改进,以获得更优的遗憾界?
- RQ5是否存在一种Thompson采样的变体,可避免在小损失情形下出现$\Omega(d\overline{L}^*)$的遗憾?
主要发现
- 本文证明了Thompson采样在半-bandit设置下可实现$\widetilde{O}(\sqrt{dL^*})$的遗憾界,与目前已知的最佳频率学遗憾界一致。
- 对尺度敏感的信息比通过捕捉与损失相关的缩放特性,实现了更紧密的首阶遗憾分析。
- 坐标熵取代了全局香农熵,首次实现了组合设置下Thompson采样在最坏情况下的最优遗憾界。
- 阈值化Thompson采样在$L^* \leq \overline{L}^*$时实现了与$T$无关的遗憾界,而标准Thompson采样不具备此性质。
- 本文证明了在上下文bandit问题中,即使$d=O(\sqrt{T})$,当$L^*=0$时,标准Thompson采样仍以高概率产生$\Omega(\sqrt{T})$的遗憾,表明其在小损失情形下失效。
- 本文首次建立了贝叶斯代理与频率学置信区间之间的新颖联系,从而构建了统一的分析框架。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。