[Paper Review] A Simple Reduction Scheme for Constrained Contextual Bandits with Adversarial Contexts via Regression
The paper introduces a modular reduction scheme that converts constrained contextual bandits with adversarial contexts into an unconstrained contextual bandit problem using online regression oracles, enabling simultaneous control of regret and constraint violations under a continuing setting.
We study constrained contextual bandits (CCB) with adversarially chosen contexts, where each action yields a random reward and incurs a random cost. We adopt the standard realizability assumption: conditioned on the observed context, rewards and costs are drawn independently from fixed distributions whose expectations belong to known function classes. We consider the continuing setting, in which the algorithm operates over the entire horizon even after the budget is exhausted. In this setting, the objective is to simultaneously control regret and cumulative constraint violation. Building on the seminal SquareCB framework of Foster et al. (2018), we propose a simple and modular algorithmic scheme that leverages online regression oracles to reduce the constrained problem to a standard unconstrained contextual bandit problem with adaptively defined surrogate reward functions. In contrast to most prior work on CCB, which focuses on stochastic contexts, our reduction yields improved guarantees for the more general adversarial context setting, together with a compact and transparent analysis.
Motivation & Objective
- Motivate and address constrained contextual bandits (CCB) with adversarially chosen contexts.
- Develop a simple, modular algorithmic scheme leveraging online regression oracles.
- Reduce CCB to an unconstrained contextual bandit with surrogate rewards.
- Provide a compact, transparent analysis under the continuing setting where budgets may be exhausted but learning continues.
Proposed method
- Build on the SquareCB framework to design a reduction scheme.
- Use online regression oracles to construct surrogate reward functions.
- Reduce the constrained problem to a standard unconstrained CB problem with these surrogates.
- Operate in a continuing setting with budget being exhausted yet learning continues.
- Provide a modular and transparent analysis for adversarial contexts.
Experimental results
Research questions
- RQ1Can constrained contextual bandits with adversarial contexts be reduced to an unconstrained CB problem using regression oracles?
- RQ2What are the regret and constraint violation guarantees under the continuing setting with this reduction?
- RQ3How does the proposed reduction compare to prior CCB approaches that focus on stochastic contexts?
- RQ4What is the impact of via surrogate rewards on performance guarantees for adversarial contexts?
Key findings
- A simple modular reduction scheme yields improved guarantees in the adversarial-context setting.
- The reduction leverages online regression oracles to define surrogate rewards for the unconstrained CB.
- The approach provides a compact and transparent analysis in the continuing setting.
- The scheme extends the SquareCB framework to constrained, adversarial-context scenarios.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.