Skip to main content
QUICK REVIEW

[Paper Review] On the Complexity of Adversarial Decision Making

Dylan J. Foster, Alexander Rakhlin|arXiv (Cornell University)|Jun 27, 2022
Advanced Bandit Algorithms Research4 citations
TL;DR

This paper establishes that the Decision-Estimation Coefficient (DEC), when applied to the convex hull of a model class, is both necessary and sufficient for achieving low regret in adversarial decision making—encompassing structured bandits and reinforcement learning with adversarial dynamics. The key result shows that convexification of the model class governs the statistical price of robustness in adversarial settings.

ABSTRACT

A central problem in online learning and decision making -- from bandits to reinforcement learning -- is to understand what modeling assumptions lead to sample-efficient learning guarantees. We consider a general adversarial decision making framework that encompasses (structured) bandit problems with adversarial rewards and reinforcement learning problems with adversarial dynamics. Our main result is to show -- via new upper and lower bounds -- that the Decision-Estimation Coefficient, a complexity measure introduced by Foster et al. in the stochastic counterpart to our setting, is necessary and sufficient to obtain low regret for adversarial decision making. However, compared to the stochastic setting, one must apply the Decision-Estimation Coefficient to the convex hull of the class of models (or, hypotheses) under consideration. This establishes that the price of accommodating adversarial rewards or dynamics is governed by the behavior of the model class under convexification, and recovers a number of existing results -- both positive and negative. En route to obtaining these guarantees, we provide new structural results that connect the Decision-Estimation Coefficient to variants of other well-known complexity measures, including the Information Ratio of Russo and Van Roy and the Exploration-by-Optimization objective of Lattimore and György.

Motivation & Objective

  • To identify the structural conditions under which sample-efficient learning is possible in adversarial decision making settings.
  • To understand the statistical complexity of online learning under adversarial rewards and dynamics, particularly in reinforcement learning and structured bandits.
  • To bridge the gap between stochastic and adversarial decision making by characterizing the necessary and sufficient conditions for regret minimization.
  • To establish a formal connection between the DEC and other complexity measures such as the Information Ratio and Exploration-by-Optimization.
  • To provide tight upper and lower bounds on regret that depend on the convex hull of the model class, revealing the price of robustness in adversarial environments.

Proposed method

  • Introduces an adversarial variant of the Decision Making with Structured Observations (DMSO) framework, where models are chosen adaptively by an adversary.
  • Defines the Decision-Estimation Coefficient (DEC) for adversarial settings and proves that its convexified version governs regret bounds.
  • Applies new structural analysis to show that the DEC in the convex hull of the model class is both necessary and sufficient for low regret.
  • Establishes connections between the convexified DEC and existing complexity measures: the Information Ratio (Russo & Van Roy, 2018) and the Exploration-by-Optimization objective (Lattimore & György, 2021).
  • Uses a reduction-based argument to derive lower bounds on regret by constructing a family of MDPs and analyzing their behavior under convexification.
  • Employs a minimax argument and a carefully designed perturbation of MDPs to derive a lower bound on the DEC that scales with the number of states and actions.

Experimental results

Research questions

  • RQ1What structural properties of a model class determine the minimax regret in adversarial decision making?
  • RQ2Is the Decision-Estimation Coefficient (DEC) sufficient and necessary for achieving low regret in adversarial settings?
  • RQ3How does convexification of the model class affect the complexity of adversarial decision making?
  • RQ4What is the relationship between the DEC and other known complexity measures such as the Information Ratio and Exploration-by-Optimization?
  • RQ5Can tight regret bounds be derived for adversarial reinforcement learning and structured bandits using the convexified DEC?

Key findings

  • The convexified Decision-Estimation Coefficient (DEC) is both necessary and sufficient for achieving low regret in adversarial decision making.
  • For any algorithm with reasonable tail behavior, the optimal regret is lower bounded by a localized version of the convexified DEC.
  • The paper establishes a regret upper bound that scales with the convexified DEC, proving its sufficiency.
  • The lower bound construction shows that the DEC in the convex hull of the model class is the fundamental complexity measure, with a lower bound of $ \frac{A^{ ext{min}\{S-1,H,K\}}}{24\gamma} $ for tabular MDPs.
  • The results recover and unify a range of existing positive and negative results in adversarial bandits and reinforcement learning.
  • The analysis reveals that the price of robustness in adversarial settings is governed by the behavior of the model class under convexification.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.