Skip to main content
QUICK REVIEW

[论文解读] Taking a hint: How to leverage loss predictors in contextual bandits?

Chen-Yu Wei, Haipeng Luo|arXiv (Cornell University)|Mar 4, 2020
Advanced Bandit Algorithms Research参考文献 29被引用 5
一句话总结

本文研究带有损失预测器的上下文Bandits问题,表明当总预测误差 𝔈 较小时,遗憾可显著降低。该文提出了新颖的算法,利用预测器,在已知 𝔈 时实现最优遗憾界 𝒪(√𝔼𝑇¹/⁴),在未知 𝔈 时实现 𝒪(√𝔼𝑇¹/³),并揭示了多个预测器会导致与 M 呈线性依赖关系,这与非上下文设置下的情况不同。

ABSTRACT

We initiate the study of learning in contextual bandits with the help of loss predictors. The main question we address is whether one can improve over the minimax regret $\mathcal{O}(\sqrt{T})$ for learning over $T$ rounds, when the total error of the predictor $\mathcal{E} \leq T$ is relatively small. We provide a complete answer to this question, including upper and lower bounds for various settings: adversarial versus stochastic environments, known versus unknown $\mathcal{E}$, and single versus multiple predictors. We show several surprising results, such as 1) the optimal regret is $\mathcal{O}(\min\{\sqrt{T}, \sqrt{\mathcal{E}}T^\frac{1}{4}\})$ when $\mathcal{E}$ is known, a sharp contrast to the standard and better bound $\mathcal{O}(\sqrt{\mathcal{E}})$ for non-contextual problems (such as multi-armed bandits); 2) the same bound cannot be achieved if $\mathcal{E}$ is unknown, but as a remedy, $\mathcal{O}(\sqrt{\mathcal{E}}T^\frac{1}{3})$ is achievable; 3) with $M$ predictors, a linear dependence on $M$ is necessary, even if logarithmic dependence is possible for non-contextual problems. We also develop several novel algorithmic techniques to achieve matching upper bounds, including 1) a key action remapping technique for optimal regret with known $\mathcal{E}$, 2) implementing Catoni's robust mean estimator efficiently via an ERM oracle leading to an efficient algorithm in the stochastic setting with optimal regret, 3) constructing an underestimator for $\mathcal{E}$ via estimating the histogram with bins of exponentially increasing size for the stochastic setting with unknown $\mathcal{E}$, and 4) a self-referential scheme for learning with multiple predictors, all of which might be of independent interest.

研究动机与目标

  • 探究损失预测器是否能在上下文Bandits中实现超越标准 𝒪(√T) 最小最大遗憾边界的遗憾改进。
  • 确定当总预测误差 𝔈 已知或未知时,最优遗憾界为何。
  • 分析多个预测器对遗憾的影响,特别是其与预测器数量 M 的依赖关系。
  • 设计在对抗性和随机设置下均能实现这些改进遗憾界的高效算法。
  • 理解上下文Bandits与非上下文问题(如多臂老虎机)在利用损失预测器方面的根本差异。

提出的方法

  • 提出一种新颖的动作重映射技术,在已知 𝔈 时实现最优遗憾,从而更好地利用预测质量。
  • 在随机设置中,通过ERM oracle实现Catoni的鲁棒均值估计器的计算高效实现,以达到最优遗憾。
  • 提出一种基于指数增长区间大小的直方图估计的 𝔈 低估器,以在随机环境中自适应处理未知 𝔈。
  • 设计一种自指机制,用于学习多个预测器,实现对 M 个来源预测的自适应聚合。
  • 采用基于预算的机制,动态调整探索程度,以在预测准确性和鲁棒性之间取得平衡。
  • 通过分层分析预测误差和置信区间,在对抗性和i.i.d.设置下界定遗憾。

实验结果

研究问题

  • RQ1损失预测器是否能在上下文Bandits中实现超越标准 𝒪(√T) 最小最大遗憾边界的遗憾降低?
  • RQ2当总预测误差 𝔈 已知时,最优遗憾界是什么?其与非上下文设置下的表现相比如何?
  • RQ3当 𝔈 未知时,遗憾界会发生什么变化?是否仍可实现次线性遗憾?
  • RQ4预测器数量 M 如何影响遗憾?在上下文Bandits中,能否实现对 M 的对数依赖?
  • RQ5能否设计出计算开销极小的高效算法,以实现最优遗憾?

主要发现

  • 在已知 𝔈 时,最优遗憾为 𝒪(min{√T, √𝔼𝑇¹/⁴}),其严格劣于非上下文问题(如多臂老虎机)中的 𝒪(√𝔼) 边界。
  • 当 𝔈 未知时,无法实现 𝒪(√𝔼𝑇¹/⁴) 边界;相反,可达到的最佳遗憾为 𝒪(√𝔼𝑇¹/³),只要 𝔈 = o(T¹/³),该边界仍为次线性。
  • 对于 M 个预测器,已知 𝔈* 时最优遗憾为 𝒪(√(M𝔼*)𝑇¹/⁴),表明其与 M 呈线性依赖,这与非上下文设置中可能实现的对数依赖不同。
  • 当 𝔈* 未知时,最优遗憾为 𝒪(M²/³(𝔼*T)¹/³),显示出预测器数量与预测误差之间的非平凡权衡。
  • 所提出的算法在所有设置下均实现了匹配的上界,采用了诸如动作重映射和基于直方图的 𝔈 鲁棒低估等新方法。
  • 研究结果建立了紧致的下界,证明在给定假设下,所推导的遗憾边界在信息论上是最优的。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。