[论文解读] Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
本文提出了一种基于分歧的框架,用于表征上下文Bandits和函数逼近强化学习中的实例相关遗憾。它提出了Oracle高效的算法,能够在可能的情况下自适应奖励差距,实现最优的差距相关样本复杂度,同时在最坏情况下保持极小极大率。
In the classical multi-armed bandit problem, instance-dependent algorithms attain improved performance on "easy" problems with a gap between the best and second-best arm. Are similar guarantees possible for contextual bandits? While positive results are known for certain special cases, there is no general theory characterizing when and how instance-dependent regret bounds for contextual bandits can be achieved for rich, general classes of policies. We introduce a family of complexity measures that are both sufficient and necessary to obtain instance-dependent regret bounds. We then introduce new oracle-efficient algorithms which adapt to the gap whenever possible, while also attaining the minimax rate in the worst case. Finally, we provide structural results that tie together a number of complexity measures previously proposed throughout contextual bandits, reinforcement learning, and active learning and elucidate their role in determining the optimal instance-dependent regret. In a large-scale empirical evaluation, we find that our approach often gives superior results for challenging exploration problems. Turning our focus to reinforcement learning with function approximation, we develop new oracle-efficient algorithms for reinforcement learning with rich observations that obtain optimal gap-dependent sample complexity.
研究动机与目标
- 解决在特殊情形之外,上下文Bandits中实例相关遗憾界缺乏一般理论的问题。
- 开发既必要又充分的复杂度度量,以实现实例相关遗憾保证。
- 设计能够自适应奖励差距、同时在最坏情况设置下保持极小极大最优性的Oracle高效算法。
- 统一并阐明各类复杂度度量(如可消解维数和星数)在上下文Bandits、强化学习和主动学习中的作用。
- 将该框架扩展至具有丰富观测的强化学习,提供最优的差距相关样本复杂度。
提出的方法
- 引入分歧系数,以衡量策略类在与奖励差距相关联的统计容量。
- 提出一类新的复杂度度量(如策略分歧系数、星数、可消解维数),以捕捉实例特定的困难程度。
- 设计一种Oracle高效的算法,利用回归Oracle估计值函数,并动态适应奖励差距。
- 使用集中不等式和信息论工具,建立分布无关且与尺度相关的遗憾界。
- 利用置信集和最小二乘估计,限制函数逼近和偏移误差在函数逼近强化学习中的影响。
- 通过并集界和同质性论证,将结果从函数类扩展到其星包络,以提升泛化性能。
实验结果
研究问题
- RQ1在上下文Bandits中,哪些复杂度度量既必要又充分以实现实例相关遗憾界?
- RQ2能否设计出能够自适应奖励差距、同时保持极小极大最优性的Oracle高效算法?
- RQ3在实例相关学习背景下,现有复杂度度量(如可消解维数和星数)之间有何关联?
- RQ4在具有丰富观测和函数逼近的强化学习中,最优的差距相关样本复杂度是什么?
- RQ5基于分歧的分析能否统一上下文Bandits、强化学习和主动学习中的理论保证?
主要发现
- 本文证明,策略分歧系数在上下文Bandits中既是实例相关遗憾界成立的必要条件,也是充分条件。
- 所提出的算法在奖励差距为正的简单情形下实现对数遗憾,同时在最坏情况下保持$\sqrt{T}$的极小极大率。
- 引入了一种新的复杂度度量——星数,其与分歧系数紧密关联,并提供了学习困难程度的结构表征。
- 在具有丰富观测的强化学习中,该算法实现了最优的差距相关样本复杂度,与已知下界一致。
- 实验评估表明,该算法在具有挑战性的探索问题上相比基线方法表现更优。
- 理论分析揭示,可消解维数和星数均为更广泛分歧框架下的特例,统一了先前的复杂度度量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。