Skip to main content
QUICK REVIEW

[论文解读] On the Complexity of Adversarial Decision Making

Dylan J. Foster, Alexander Rakhlin|arXiv (Cornell University)|Jun 27, 2022
Advanced Bandit Algorithms Research被引用 4
一句话总结

本文证明,当应用于模型类的凸包时,决策-估计系数(DEC)在对抗性决策制定中既是实现低遗憾的必要条件也是充分条件——涵盖结构化多臂赌博机和具有对抗性动态的强化学习。关键结果表明,模型类的凸化决定了对抗性环境中鲁棒性的统计代价。

ABSTRACT

A central problem in online learning and decision making -- from bandits to reinforcement learning -- is to understand what modeling assumptions lead to sample-efficient learning guarantees. We consider a general adversarial decision making framework that encompasses (structured) bandit problems with adversarial rewards and reinforcement learning problems with adversarial dynamics. Our main result is to show -- via new upper and lower bounds -- that the Decision-Estimation Coefficient, a complexity measure introduced by Foster et al. in the stochastic counterpart to our setting, is necessary and sufficient to obtain low regret for adversarial decision making. However, compared to the stochastic setting, one must apply the Decision-Estimation Coefficient to the convex hull of the class of models (or, hypotheses) under consideration. This establishes that the price of accommodating adversarial rewards or dynamics is governed by the behavior of the model class under convexification, and recovers a number of existing results -- both positive and negative. En route to obtaining these guarantees, we provide new structural results that connect the Decision-Estimation Coefficient to variants of other well-known complexity measures, including the Information Ratio of Russo and Van Roy and the Exploration-by-Optimization objective of Lattimore and György.

研究动机与目标

  • 确定在对抗性决策制定设置中实现样本高效学习的结构性条件。
  • 理解在对抗性奖励和动态下在线学习的统计复杂度,特别是在强化学习和结构化多臂赌博机中。
  • 通过刻画遗憾最小化的必要和充分条件,弥合随机与对抗性决策制定之间的差距。
  • 建立DEC与其他复杂度度量(如信息比和基于优化的探索)之间的正式联系。
  • 提供依赖于模型类凸包的紧致遗憾上下界,揭示对抗环境中鲁棒性的代价。

提出的方法

  • 引入对抗性结构观测决策制定(DMSO)框架的对抗变体,其中模型由对手自适应选择。
  • 为对抗性设置定义决策-估计系数(DEC),并证明其凸化版本控制遗憾界。
  • 应用新的结构分析表明,模型类凸包中的DEC是实现低遗憾的必要且充分条件。
  • 建立凸化DEC与现有复杂度度量之间的联系:信息比(Russo & Van Roy, 2018)和基于优化的探索目标(Lattimore & György, 2021)。
  • 通过基于归约的论证,通过构造一类MDP并分析其在凸化下的行为,推导出遗憾的下界。
  • 采用极小极大论证法,并对MDP进行精心设计的扰动,推导出与状态数和动作数成比例的DEC下界。

实验结果

研究问题

  • RQ1模型类的何种结构性质决定了对抗性决策制定中的最小最大遗憾?
  • RQ2决策-估计系数(DEC)在对抗性环境中是否足以实现低遗憾,且是否为必要条件?
  • RQ3模型类的凸化如何影响对抗性决策制定的复杂度?
  • RQ4DEC与其他已知复杂度度量(如信息比和基于优化的探索)之间存在何种关系?
  • RQ5能否使用凸化的DEC为对抗性强化学习和结构化多臂赌博机推导出紧致的遗憾界?

主要发现

  • 凸化的决策-估计系数(DEC)在对抗性决策制定中既是实现低遗憾的必要条件也是充分条件。
  • 对于具有合理尾部行为的任何算法,最优遗憾均被局部化的凸化DEC下界所限制。
  • 本文建立了与凸化DEC成比例的遗憾上界,证明了其充分性。
  • 下界构造表明,模型类凸包中的DEC是根本的复杂度度量,对于表格型MDP,其下界为$ \frac{A^{\text{min}\{S-1,H,K\}}}{24\gamma} $。
  • 结果统一并恢复了对抗性多臂赌博机和强化学习中一系列现有的正负结果。
  • 分析表明,对抗性环境中鲁棒性的代价由模型类在凸化下的行为所决定。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。