Skip to main content
QUICK REVIEW

[论文解读] Strategy-Driven Limit Theorems Associated Bandit Problems

Zengjing Chen, Shui Feng|arXiv (Cornell University)|Apr 9, 2022
Advanced Bandit Algorithms Research被引用 4
一句话总结

本文提出了针对两臂老虎机问题的策略驱动极限定理,建立了依赖于采样策略的弱大数定律、大偏差原理和中心极限定理。关键贡献在于识别出依赖于策略的极限分布——非正态且与集合相关的分布——揭示了学习结构,并实现了对最大/最小奖励的估计,同时避免了帕伦多悖论。

ABSTRACT

Motivated by the study of asymptotic behaviour of the bandit problems, we obtain several strategy-driven limit theorems including the law of large numbers, the large deviation principle, and the central limit theorem. Different from the classical limit theorems, we develop sampling strategy-driven limit theorems that generate the maximum or minimum average reward. The law of large numbers identifies all possible limits that are achievable under various strategies. The large deviation principle provides the maximum decay probabilities for deviations from the limiting domain. To describe the fluctuations around averages, we obtain strategy-driven central limit theorems under optimal strategies. The limits in these theorem are identified explicitly, and depend heavily on the structure of the events or the integrating functions and strategies. This demonstrates the key signature of the learning structure. Our results can be used to estimate the maximal (minimal) rewards, and to identify the conditions of avoiding the Parrondo's paradox in the two-armed bandit problem. It also lays the theoretical foundation for statistical inference in determining the arm that offers the higher mean reward.

研究动机与目标

  • 开发一个策略驱动的极限定理框架,以表征两臂老虎机问题中采样策略的渐近行为。
  • 识别在各种采样策略下平均奖励的所有可实现极限,扩展经典的大数定律。
  • 提供一个大偏差原理,量化偏离最优奖励极限的偏差概率的最大衰减速率。
  • 推导一个具有非正态极限分布的战略中心极限定理,其分布依赖于采样策略和积分函数的结构。
  • 将该框架应用于估计最大/最小奖励,并确定在老虎机设置中避免帕伦多悖论的条件。

提出的方法

  • 提出一种策略驱动的弱大数定律和强大数定律,通过引入采样策略依赖性,推广了罗宾斯的结果。
  • 建立一个大偏差原理(LDP),其速率函数为 $ I(x) = \inf_{\alpha \in [0,1]} I_\alpha(x) $,其中 $ I_\alpha(x) $ 是在策略 $ \theta^\alpha $ 下混合矩生成函数的导出结果。
  • 提出一种战略中心极限定理,其极限分布被明确识别为非正态且与集合相关,取决于策略和奖励结构。
  • 利用克拉默定理和压缩原理,从策略序列 $ \theta^\alpha $ 下的奖励经验测度推导出LDP的速率函数。
  • 构建一族策略 $ \theta^\alpha $,以控制每条臂在长期中被拉动的比例,从而推导出依赖于策略的极限。
  • 采用非线性概率和勒让德-芬赫尔变换来表征速率函数 $ \Lambda^* $,将矩生成函数与大偏差速率联系起来。

实验结果

研究问题

  • RQ1在两臂老虎机问题中,不同采样策略下可实现的平均奖励所有可能极限的集合是什么?
  • RQ2在任意策略下,偏离最优奖励极限的偏差概率的最大衰减速率是多少?
  • RQ3在最优采样策略下,平均奖励周围的波动行为如何?其极限分布的形式是什么?
  • RQ4该框架能否用于估计在未知臂均值下的序贯采样中最大或最小期望奖励?
  • RQ5在使用这种策略驱动方法时,两臂老虎机设置中帕伦多悖论在何种条件下可以避免?

主要发现

  • 策略驱动的大数定律识别出平均奖励的所有可能极限,即经验均值的极限,其值由策略和基础奖励分布明确决定。
  • 大偏差原理提供了速率函数 $ I(x) = \inf_{\alpha \in [0,1]} I_\alpha(x) $,其中 $ I_\alpha(x) $ 是在策略 $ \theta^\alpha $ 下从矩生成函数混合导出的结果,并表明在最优均值处衰减速率达到最小。
  • 战略中心极限定理得出的极限分布通常为非正态,且依赖于策略和积分函数的结构,其显式形式为 $ I_\alpha(x) = \inf\{ \alpha \Lambda^*_{\mu_L}(y) + (1-\alpha) \Lambda^*_{\mu_R}(z) : \alpha y + (1-\alpha) z = x \} $。
  • 该框架可通过识别实现极端极限的最优策略,实现对最大和最小期望奖励的估计。
  • 结果为在最优采样策略下进行假设检验的统计推断提供了理论基础,以识别具有更高均值的臂。
  • 分析表明,当仔细分析策略依赖的极限和速率函数时,帕伦多悖论可以被避免,特别是当最优策略不会从两个亏损组件中导致悖论性胜利时。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。