Skip to main content
QUICK REVIEW

[论文解读] Diffusion Approximations for Thompson Sampling in the Small Gap Regime

Lin Fan, Peter W. Glynn|arXiv (Cornell University)|May 19, 2021
Advanced Bandit Algorithms Research参考文献 53被引用 7
一句话总结

本文在小间隙区间内建立了 Thompson 采样弱收敛理论,其中臂均值差异按 $1/\sqrt{n}$ 缩放,表明该算法的动力学行为弱收敛于随机微分方程(SDE)和随机常微分方程(random ODE)的解。其核心贡献是一个不变性原理,证明了在不同奖励分布和先验下,扩散极限是普适的,且与正态分布情况下的极限相同。

ABSTRACT

We study the process-level dynamics of Thompson sampling in the ``small gap'' regime. The small gap regime is one in which the gaps between the arm means are of order $\sqrtγ$ or smaller and the time horizon is of order $1/γ$, where $γ$ is small. As $γ\downarrow 0$, we show that the process-level dynamics of Thompson sampling converge weakly to the solutions to certain stochastic differential equations and stochastic ordinary differential equations. Our weak convergence theory is developed from first principles using the Continuous Mapping Theorem, can handle stationary, weakly dependent reward processes, and can also be adapted to analyze a variety of sampling-based bandit algorithms. Indeed, we show that the process-level dynamics of many sampling-based bandit algorithms -- including Thompson sampling designed for any single-parameter exponential family of rewards, as well as non-parametric bandit algorithms based on bootstrap re-sampling -- satisfy an invariance principle. Namely, their weak limits coincide with that of Gaussian parametric Thompson sampling with Gaussian priors. Moreover, in the small gap regime, the regret performance of these algorithms is generally insensitive to model mis-specification, changing continuously with increasing degrees of mis-specification.

研究动机与目标

  • 为在小间隙区间内分析 Thompson 采样建立严格的弱收敛框架。
  • 在臂间隙按 $\Delta/\sqrt{n}$ 缩放下,刻画 Thompson 采样动力学在时间范围 $n \to \infty$ 时的渐近行为。
  • 证明在 $1/\sqrt{n}$ 缩放下,极限扩散过程对不同的奖励分布和先验是普适的。
  • 将弱收敛理论扩展至多臂和线性 bandit 设置,包括时变动作集。
  • 使用连续映射定理从第一原理出发推导,使该理论可推广至其他基于采样的 bandit 算法。

提出的方法

  • 使用弱收敛(依分布收敛)分析在臂间隙按 $\Delta/\sqrt{n}$ 缩放下的 Thompson 采样动力学。
  • 推导出在有限 $n$ 下控制 Thompson 采样行为的 SDE 和随机 ODE 的离散时间类比。
  • 应用连续映射定理和 $D_{\mathbb{R}^d}[0,1]$ 中的紧性准则,证明弱收敛至连续扩散极限。
  • 利用 Glivenko-Cantelli 定理和 bracketing entropy 条件,验证后验分位数函数的统计过程收敛性。
  • 使用鞅函数中心极限定理,建立收敛至具有特定协方差结构的布朗运动。
  • 通过矩界和增量方差控制,应用 $D_{\mathbb{R}}[0,1]$ 中紧性的充分条件。

实验结果

研究问题

  • RQ1当臂均值之间的差距按 $1/\sqrt{n}$ 缩放时,Thompson 采样的动力学行为在渐近下如何表现?
  • RQ2Thompson 采样的离散动力学在分布上收敛至何种连续时间随机过程?
  • RQ3Thompson 采样的弱扩散极限是否对不同的奖励分布和先验选择具有普适性?
  • RQ4该弱收敛框架能否扩展至具有有限或无限动作集的线性 bandit,包括时变动作集?
  • RQ5后验近似和自助采样在塑造极限扩散过程中的作用是什么?

主要发现

  • 在 $\Delta/\sqrt{n}$ 缩放下,当 $n \to \infty$ 时,Thompson 采样动力学弱收敛于 SDE 和随机 ODE 的解。
  • 极限扩散过程具有普适性:只要后验被良好近似,其极限与具体奖励分布或先验无关。
  • 弱扩散极限与正态分布奖励和先验情况下的极限一致,证实了经典的 Bernstein-von Mises 近似。
  • 存在一个不变性原理:对于使用后验近似或自助采样的 Thompson 采样及相关算法,在 $1/\sqrt{n}$ 间隙区间下,其弱极限是相同的。
  • 该收敛框架具有通用性,可直接用于分析其他基于采样的 bandit 算法,如基于自助采样的探索策略。
  • 该理论适用于多臂 bandit 和具有有限或无限动作集的线性 bandit,包括随机和时变动作集。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。