Skip to main content
QUICK REVIEW

[論文レビュー] A Simple Reduction Scheme for Constrained Contextual Bandits with Adversarial Contexts via Regression

Dhruv Sarkar, Abhishek Sinha|arXiv (Cornell University)|Feb 4, 2026
Advanced Bandit Algorithms Research被引用数 0
ひとこと要約

本論文は、制約付き文脈 bandits を敵対的文脈でオンライン回帰オラクルを用いて無制約の文脈 bandit 問題へ変換するモジュラー削減スキームを提案し、継続設定における後悔と制約違反の同時制御を可能にする。

ABSTRACT

We study constrained contextual bandits (CCB) with adversarially chosen contexts, where each action yields a random reward and incurs a random cost. We adopt the standard realizability assumption: conditioned on the observed context, rewards and costs are drawn independently from fixed distributions whose expectations belong to known function classes. We consider the continuing setting, in which the algorithm operates over the entire horizon even after the budget is exhausted. In this setting, the objective is to simultaneously control regret and cumulative constraint violation. Building on the seminal SquareCB framework of Foster et al. (2018), we propose a simple and modular algorithmic scheme that leverages online regression oracles to reduce the constrained problem to a standard unconstrained contextual bandit problem with adaptively defined surrogate reward functions. In contrast to most prior work on CCB, which focuses on stochastic contexts, our reduction yields improved guarantees for the more general adversarial context setting, together with a compact and transparent analysis.

研究の動機と目的

  • 敵対的に選択された文脈を伴う制約付き文脈 bandits (CCB) の動機付けと課題設定。
  • オンライン回帰オラクルを活用した簡潔でモジュラーなアルゴリズムスキームの開発。
  • surrogate reward を用いて CCB を無制約の文脈 bandit 問題へ還元。
  • 予算が尽きても学習が継続する継続設定の下で、コンパクトで透明な分析を提供。

提案手法

  • SquareCB フレームワークを土台に削減スキームを設計。
  • オンライン回帰オラクルを用いて surrogate reward 関数を構築。
  • これらの surrogate を用いて制約付き問題を標準的な無制約 CB 問題へ還元。
  • 予算が尽きても学習が継続する継続設定で運用。
  • 敵対的文脈に対してモジュールで透明な分析を提供。

実験結果

リサーチクエスチョン

  • RQ1敵対的文脈を伴う制約付き文脈 bandits を回帰オラクルを用いて無制約 CB 問題へ還元できるか。
  • RQ2この削減による継続設定での後悔と制約違反の保証はどうなるか。
  • RQ3確率的文脈に焦点を当てた従来の CCB アプローチとこの提案削減の比較はどうなるか。
  • RQ4 surrogate rewards が敵対的文脈の性能保証に与える影響はどの程度か。

主な発見

  • 敵対的文脈設定で改善された保証をもたらす簡易なモジュラー削減スキーム。
  • オンライン回帰オラクルを活用して無制約 CB の surrogate rewards を定義。
  • 継続設定でのコンパクトで透明な分析を提供。
  • 本手法は SquareCB フレームワークを制約付き・敵対的文脈シナリオへ拡張。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。