Skip to main content
QUICK REVIEW

[论文解读] Tracking the Best Expert in Non-stationary Stochastic Environments

Chen-Yu Wei, Yi-Te Hong|arXiv (Cornell University)|Dec 2, 2017
Advanced Bandit Algorithms Research参考文献 13被引用 16
一句话总结

本论文引入了一个新参数 $Λ$,用于衡量非平稳随机环境中损失分布的总统计方差,以改进多臂赌博机和完整信息设置下的遗憾分析。研究证明,即使 $Λ$ 保持恒定,带 bandit 的遗憾仍随 $T$ 增长;而在完整信息设置下,当 $Γ$ 和 $Λ$ 恒定时,遗憾可为常数;当 $V$ 和 $Λ$ 恒定时,遗憾呈现 $T^{1/3}$ 依赖关系,且提供了匹配的上下界。

ABSTRACT

We study the dynamic regret of multi-armed bandit and experts problem in non-stationary stochastic environments. We introduce a new parameter $Λ$, which measures the total statistical variance of the loss distributions over $T$ rounds of the process, and study how this amount affects the regret. We investigate the interaction between $Λ$ and $Γ$, which counts the number of times the distributions change, as well as $Λ$ and $V$, which measures how far the distributions deviates over time. One striking result we find is that even when $Γ$, $V$, and $Λ$ are all restricted to constant, the regret lower bound in the bandit setting still grows with $T$. The other highlight is that in the full-information setting, a constant regret becomes achievable with constant $Γ$ and $Λ$, as it can be made independent of $T$, while with constant $V$ and $Λ$, the regret still has a $T^{1/3}$ dependency. We not only propose algorithms with upper bound guarantee, but prove their matching lower bounds as well.

研究动机与目标

  • 通过引入一个衡量损失分布总统计方差的新参数 $Λ$,理解非平稳性如何影响随机在线学习中的遗憾。
  • 分析 $Λ$、$Γ$(分布变化次数)和 $V$(均值总偏移量)在决定遗憾界方面的相互作用。
  • 通过证明匹配的遗憾下界,弥合现有上界与下界在非平稳随机设置下的差距。

提出的方法

  • 将 $Λ$ 定义为 $T$ 轮中损失分布方差之和,提供一种超越 $Γ$ 和 $V$ 的非平稳性精细化度量。
  • 采用基于归约的证明策略,通过构造迫使任何算法承受高遗憾的对抗性损失分布序列,推导遗憾下界。
  • 应用条件期望论证,通过 KL 散度和集中不等式,界定向时间区间内算法选择次优臂的期望次数。
  • 采用递归构造方法,在长度为 $B = \sqrt[3]{\Lambda T / (32K V^2)}$ 的区间内构建损失分布,确保方差和均值漂移受控。
  • 利用引理 4.4 保证存在均值差距 $\epsilon$、方差 $\sigma^2$ 和 KL 散度 $\leq \epsilon^2 / \sigma^2$ 的分布 $\mathcal{P}$ 和 $\mathcal{Q}$,从而实现紧密的遗憾分析。
  • 通过分析 $\Lambda$、$V$、$\Gamma$ 和 $T$ 之间的权衡,在赌博机和完整信息设置中均推导出匹配的上下界。

实验结果

研究问题

  • RQ1损失分布的总统计方差 $\Lambda$ 如何影响非平稳随机环境中的遗憾?
  • RQ2当 $\Gamma$ 和 $\Lambda$ 有界时,即使 $T$ 增大,是否可在完整信息设置中实现常数遗憾?
  • RQ3当 $\Gamma$、$V$ 和 $\Lambda$ 均为常数时,赌博机设置中的遗憾基本极限是什么?
  • RQ4$\Lambda$、$V$ 和 $\Gamma$ 如何相互作用,以决定非平稳随机在线学习中的遗憾下界?
  • RQ5当 $V$ 和 $\Lambda$ 恒定时,即使在完整信息设置中,遗憾的 $T^{1/3}$ 依赖关系是否不可避免?

主要发现

  • 在赌博机设置中,即使 $\Gamma$、$V$ 和 $\Lambda$ 恒定,遗憾下界仍以 $\Omega(\sqrt[3]{\Lambda V T} + \sqrt{V T})$ 的形式随 $T$ 增长,表明尽管非平稳性受控,$T$-依赖性依然存在。
  • 在完整信息设置中,当 $\Gamma$ 和 $\Lambda$ 均有界时,可实现常数遗憾,使遗憾与 $T$ 无关。
  • 当 $V$ 和 $\Lambda$ 恒定时,完整信息设置中的遗憾下界为 $\Omega(\sqrt[3]{\Lambda V T})$,与上界匹配,表明 $T^{1/3}$ 依赖性不可避免。
  • 本论文证明 $\sqrt[3]{\Lambda V T}$ 项在完整信息和赌博机设置中均是紧的,上下界匹配。
  • 通过引理 4.4 明确证明了存在具有受控均值差距、方差和 KL 散度的分布 $\mathcal{P}$ 和 $\mathcal{Q}$,从而支持了困难实例的构造以获得遗憾下界。
  • 分析表明,赌博机设置在本质上比完整信息设置更具挑战性,因为即使 $\Lambda$、$\Gamma$ 和 $V$ 有界,也无法消除遗憾中的 $T$-依赖性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。