Skip to main content
QUICK REVIEW

[论文解读] Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective

Dylan J. Foster, Alexander Rakhlin|arXiv (Cornell University)|Oct 7, 2020
Advanced Bandit Algorithms Research参考文献 73被引用 21
一句话总结

本文提出了一种基于分歧的框架,用于表征上下文Bandits和函数逼近强化学习中的实例相关遗憾。它提出了Oracle高效的算法,能够在可能的情况下自适应奖励差距,实现最优的差距相关样本复杂度,同时在最坏情况下保持极小极大率。

ABSTRACT

In the classical multi-armed bandit problem, instance-dependent algorithms attain improved performance on "easy" problems with a gap between the best and second-best arm. Are similar guarantees possible for contextual bandits? While positive results are known for certain special cases, there is no general theory characterizing when and how instance-dependent regret bounds for contextual bandits can be achieved for rich, general classes of policies. We introduce a family of complexity measures that are both sufficient and necessary to obtain instance-dependent regret bounds. We then introduce new oracle-efficient algorithms which adapt to the gap whenever possible, while also attaining the minimax rate in the worst case. Finally, we provide structural results that tie together a number of complexity measures previously proposed throughout contextual bandits, reinforcement learning, and active learning and elucidate their role in determining the optimal instance-dependent regret. In a large-scale empirical evaluation, we find that our approach often gives superior results for challenging exploration problems. Turning our focus to reinforcement learning with function approximation, we develop new oracle-efficient algorithms for reinforcement learning with rich observations that obtain optimal gap-dependent sample complexity.

研究动机与目标

  • 解决在特殊情形之外,上下文Bandits中实例相关遗憾界缺乏一般理论的问题。
  • 开发既必要又充分的复杂度度量,以实现实例相关遗憾保证。
  • 设计能够自适应奖励差距、同时在最坏情况设置下保持极小极大最优性的Oracle高效算法。
  • 统一并阐明各类复杂度度量(如可消解维数和星数)在上下文Bandits、强化学习和主动学习中的作用。
  • 将该框架扩展至具有丰富观测的强化学习,提供最优的差距相关样本复杂度。

提出的方法

  • 引入分歧系数,以衡量策略类在与奖励差距相关联的统计容量。
  • 提出一类新的复杂度度量(如策略分歧系数、星数、可消解维数),以捕捉实例特定的困难程度。
  • 设计一种Oracle高效的算法,利用回归Oracle估计值函数,并动态适应奖励差距。
  • 使用集中不等式和信息论工具,建立分布无关且与尺度相关的遗憾界。
  • 利用置信集和最小二乘估计,限制函数逼近和偏移误差在函数逼近强化学习中的影响。
  • 通过并集界和同质性论证,将结果从函数类扩展到其星包络,以提升泛化性能。

实验结果

研究问题

  • RQ1在上下文Bandits中,哪些复杂度度量既必要又充分以实现实例相关遗憾界?
  • RQ2能否设计出能够自适应奖励差距、同时保持极小极大最优性的Oracle高效算法?
  • RQ3在实例相关学习背景下,现有复杂度度量(如可消解维数和星数)之间有何关联?
  • RQ4在具有丰富观测和函数逼近的强化学习中,最优的差距相关样本复杂度是什么?
  • RQ5基于分歧的分析能否统一上下文Bandits、强化学习和主动学习中的理论保证?

主要发现

  • 本文证明,策略分歧系数在上下文Bandits中既是实例相关遗憾界成立的必要条件,也是充分条件。
  • 所提出的算法在奖励差距为正的简单情形下实现对数遗憾,同时在最坏情况下保持$\sqrt{T}$的极小极大率。
  • 引入了一种新的复杂度度量——星数,其与分歧系数紧密关联,并提供了学习困难程度的结构表征。
  • 在具有丰富观测的强化学习中,该算法实现了最优的差距相关样本复杂度,与已知下界一致。
  • 实验评估表明,该算法在具有挑战性的探索问题上相比基线方法表现更优。
  • 理论分析揭示,可消解维数和星数均为更广泛分歧框架下的特例,统一了先前的复杂度度量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。