Skip to main content
QUICK REVIEW

[논문 리뷰] Truly Adapting to Adversarial Constraints in Constrained MABs

Francesco Emanuele Stradi, Kalana Kalupahana|arXiv (Cornell University)|2026. 02. 16.
Advanced Bandit Algorithms Research인용 수 0
한 줄 요약

논문은 알려지지 않은, 잠재적으로 적대적 제약과 비정상(non-stationary) 손실이 있는 제약된 다암밴딧을 연구합니다. 제약 비정상성에만 감소하는 비다항식적 regret와 제약 위반을 달성하는 알고리즘을 제시하며, 전체 피드백과 밴딧 피드백 설정에서 성능을 보장합니다.

ABSTRACT

We study the constrained variant of the \emph{multi-armed bandit} (MAB) problem, in which the learner aims not only at minimizing the total loss incurred during the learning dynamic, but also at controlling the violation of multiple \emph{unknown} constraints, under both \emph{full} and \emph{bandit feedback}. We consider a non-stationary environment that subsumes both stochastic and adversarial models and where, at each round, both losses and constraints are drawn from distributions that may change arbitrarily over time. In such a setting, it is provably not possible to guarantee both sublinear regret and sublinear violation. Accordingly, prior work has mainly focused either on settings with stochastic constraints or on relaxing the benchmark with fully adversarial constraints (\emph{e.g.}, via competitive ratios with respect to the optimum). We provide the first algorithms that achieve optimal rates of regret and \emph{positive} constraint violation when the constraints are stochastic while the losses may vary arbitrarily, and that simultaneously yield guarantees that degrade smoothly with the degree of adversariality of the constraints. Specifically, under \emph{full feedback} we propose an algorithm attaining $\widetilde{\mathcal{O}}(\sqrt{T}+C)$ regret and $\widetilde{\mathcal{O}}(\sqrt{T}+C)$ {positive} violation, where $C$ quantifies the amount of non-stationarity in the constraints. We then show how to extend these guarantees when only bandit feedback is available for the losses. Finally, when \emph{bandit feedback} is available for the constraints, we design an algorithm achieving $\widetilde{\mathcal{O}}(\sqrt{T}+C)$ {positive} violation and $\widetilde{\mathcal{O}}(\sqrt{T}+C\sqrt{T})$ regret.

연구 동기 및 목표

  • 未知하고 시간에 따라 달라지는 제약 분포가 제약된 MAB에 미치는 영향을 이해한다.
  • 손실이 적대적일 수 있지만 제약이 확률적일 때 비다항식적 regret와 비다항식적 양의 제약 위반을 달성하는 알고리즘을 개발한다.
  • 손실과 제약에 대한 밴딧 피드백으로의 확장을 제공한다.
  • 제약의 비정상성 수준 C에 따라 위반과 regret 한계가 어떻게 악화되는지 특징지운다.

제안 방법

  • 제약 비정상성을 정량화하기 위해 오염 수준 C를 도입한다.
  • 제한 위반의 낙관적 추정치를 사용하여 매 라운드에서 근사 가능한 feasible 집합 X_t를 구성한다.
  • 이동하는 의사결정 공간을 처리하고 스위칭 regret 보장을 달성하기 위해 고정 공유 업데이트가 적용된 온라인 미러 디센트를 사용한다.
  • 손실에 대한 밴딧 피드백에 대해 충분한 탐사를 보장하는 두 단계 접근법(ExpOpt-ConOMD)을 개발한다.
  • 확신 구간과 탐색 전략을 적응시켜 제약에 대한 밴딧 피드백으로 확장(Constrained OMD 변형)한다.
  • 전체 피드백 하에서 R_T = Ŝ(√T + C) 및 V_T = Ŝ(√T + C)의 이론적 경계를 보이고, 밴딧 설정으로의 확장은 비슷하거나 약간 약한 보장을 제공한다.

실험 결과

연구 질문

  • RQ1제약이 알려지지 않고 비정상적일 때 손실이 적대적일 수 있는 상황에서 비다항식적 regret와 비다항식적 양의 제약 위반을 달성할 수 있는가?
  • RQ2알 수 없는 제약 오염에 대처하면서도 합리적인 regret를 유지하기 위해 학습자가 feasible 행동 집합을 어떻게 적응적으로 구성해야 하는가?
  • RQ3전체 피드백 대 밴딧 피드백에서 손실과 제약에 대한 최적의 regret 및 위반 한계는 무엇인가?
  • RQ4제약의 비정상성 정도(C)에 따라 한계가 어떻게 악화되고, 이를 매끄럽게 악화되게 할 수 있는가?

주요 결과

  • 전체 피드백에서 제안된 알고리즘 ConOMD-FS는 regret와 양의 제약 위반이 Ŝ(√T + C) 수준이다.
  • 손실에 대해 밴딧 피드백만 이용 가능한 경우, ConOMD-FS 접근법은 분석을 조정하여 유사한 보장을 얻도록 확장된다.
  • 제약에 대한 밴딧 피드백으로, ExpOpt-ConOMD 계열은 Ŝ(√T + C) 양의 위반과 Ŝ(√T + C√T) regret를 달성하며(β = 1/2를 선택하면 이 경계가 얻어진다).
  • 오염 수준 C는 성능 저하의 주요 원동력으로 나타나며, 보장은 C에 따라 매끄럽게 저하되고 재난적으로 악화되지 않는다.
  • 제안된 방법은 의사결정 공간의 이동(시간에 따라 변하는 제약 미래)을 스위칭 regret와 페이즈 간 더블링 트릭으로 처리한다.
  • 이전 연구와 비교할 때, 결과는 확률적 설정에서 최적의 보장에 근접하며, 약한 적대적 제약 하에서도 비다항식적 regret를 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.