Skip to main content
QUICK REVIEW

[论文解读] Phase Transitions in Bandits with Switching Constraints

David Simchi‐Levi, Yunzong Xu|arXiv (Cornell University)|May 26, 2019
Advanced Bandit Algorithms Research参考文献 15被引用 4
一句话总结

本文研究了在总切换成本存在严格约束的随机多臂赌博机问题,提出了带切换约束的赌博机(BwSC)框架。该研究建立了极小化最大遗憾的匹配上下界,揭示了非传统的遗憾增长模式以及与切换预算和图结构相关的新型遗憾率相变现象,通过基于对抗TSP的索引精确推导出其依赖关系。

ABSTRACT

We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given switching budget. For this problem, we prove matching upper and lower bounds on the optimal (i.e., minimax) regret, and provide efficient rate-optimal algorithms. Surprisingly, the optimal regret of this problem exhibits a non-conventional growth rate in terms of the time horizon and the number of arms. Consequently, we discover surprising "phase transitions" regarding how the optimal regret rate changes with respect to the switching budget: when the number of arms is fixed, there are equal-length phases, where the optimal regret rate remains (almost) the same within each phase and exhibits abrupt changes between phases; when the number of arms grows with the time horizon, such abrupt changes become subtler and may disappear, but a generalized notion of phase transitions involving certain new measurements still exists. The results enable us to fully characterize the trade-off between the regret rate and the incurred switching cost in the stochastic multi-armed bandit problem, contributing new insights to this fundamental problem. Under the general switching cost structure, the results reveal interesting connections between bandit problems and graph traversal problems, such as the shortest Hamiltonian path problem.

研究动机与目标

  • 建模并分析在总切换成本存在硬性约束(而非在目标函数中惩罚切换)的随机多臂赌博机问题。
  • 刻画在线学习中遗憾与切换成本之间的基本权衡,特别是在切换成本异质且非参数化的情况下。
  • 识别并分析最优遗憾率随切换预算和臂的数量变化而产生的相变现象。
  • 为一般切换成本图(包括固定和增长的臂数)建立紧致的遗憾边界。
  • 通过一种新颖的结构索引,揭示赌博机问题与图遍历问题(尤其是对抗TSP)之间的联系。

提出的方法

  • 提出带切换约束的赌博机(BwSC)框架,其中总切换成本被作为硬性约束而非惩罚项。
  • 引入一个关键的结构索引 $ q(S,G) $,该索引源自切换图 $ G $ 上的对抗旅行商问题的动态规划公式。
  • 利用对抗TSP来建模在不确定性下最坏情况的切换序列,其中对手在每次访问后从目标集合中移除顶点。
  • 推导出最优遗憾率 $ \widetilde{\Theta}(T^{1/(2 - 2^{-q(S,G)})}) $,其指数由 $ q(S,G) $ 决定,该指数捕捉了在最坏对抗性消除规则下访问所有臂的复杂性。
  • 为固定 $ K $ 建立匹配的上下界,通过一种新颖的基于索引的表征证明其紧致性。
  • 分析当切换预算 $ S $ 增加时,遗憾率相变行为的变化,表明在固定 $ K $ 时存在突变式转变,而当 $ K $ 随 $ T $ 增长时,转变则更为平滑。

实验结果

研究问题

  • RQ1当切换受到约束而非惩罚时,最优遗憾率如何随时间范围 $ T $ 和切换预算 $ S $ 变化?
  • RQ2切换成本图 $ G $ 的哪些结构特性决定了BwSC框架中的遗憾率?
  • RQ3遗憾率的相变如何随切换预算 $ S $ 变化?其与臂的数量 $ K $ 的关系如何?
  • RQ4最优遗憾对 $ T $ 的精确依赖关系是什么?其如何随图结构 $ G $ 变化,特别是当 $ K $ 固定时?
  • RQ5BwSC问题与图遍历问题(如哈密顿路径和对抗TSP)之间存在何种联系?

主要发现

  • BwSC的最优遗憾为 $ \widetilde{\Theta}(T^{1/(2 - 2^{-q(S,G)})}) $,其中 $ q(S,G) $ 是基于切换图 $ G $ 上的对抗TSP公式推导出的结构索引。
  • 对于固定 $ K $,遗憾率表现出若干长度相等的阶段,各阶段内速率几乎恒定,且随着 $ S $ 增加,阶段间出现突变式转变。
  • 当 $ K $ 随 $ T $ 增长时,相变变得更为微妙甚至可能消失,但通过索引 $ q(S,G) $ 仍存在广义相变。
  • 对抗TSP公式为访问所有臂所需的最坏情况切换成本提供了紧致下界,其解用于定义关键索引 $ q(S,G) $。
  • 本文表明,先前工作中关于一般 $ G $ 的上下界在 $ T $ 上均不紧致,通过定理8提供了新的精确表征。
  • 索引 $ q(S,G) $ 并不总是等于其他候选索引,表明其在捕捉切换复杂性方面具有独特性和非平凡结构。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。