Skip to main content
QUICK REVIEW

[Paper Review] Finite-time Regret Bound of a Bandit Algorithm for the Semi-bounded Support Model

Junya Honda, Akimichi Takemura|arXiv (Cornell University)|Feb 10, 2012
Advanced Bandit Algorithms Research9 references3 citations
TL;DR

This paper establishes the first finite-time regret bound for the DMED bandit algorithm in the semi-bounded support model, where rewards are supported on $(-\infty,1]$ and have a moment generating function. By refining large deviation probabilities and relaxing the need for a lower support bound, the authors prove that DMED achieves optimal asymptotic performance with a non-asymptotic regret bound, extending its applicability beyond bounded $[0,1]$-supported rewards.

ABSTRACT

In this paper we consider stochastic multiarmed bandit problems. Recently a policy, DMED, is proposed and proved to achieve the asymptotic bound for the model that each reward distribution is supported in a known bounded interval, e.g. [0,1]. However, the derived regret bound is described in an asymptotic form and the performance in finite time has been unknown. We inspect this policy and derive a finite-time regret bound by refining large deviation probabilities to a simple finite form. Further, this observation reveals that the assumption on the lower-boundedness of the support is not essential and can be replaced with a weaker one, the existence of the moment generating function.

Motivation & Objective

  • To close the gap between asymptotic optimality and finite-time performance of the DMED bandit algorithm in stochastic multi-armed bandit problems.
  • To relax the assumption of a known lower support bound (e.g., 0 in $[0,1]$) by showing it is not essential for optimal regret.
  • To derive a non-asymptotic regret bound for DMED under the weaker condition that the reward distribution has a moment generating function in a neighborhood of zero.
  • To demonstrate that the theoretical lower bound on regret remains unchanged when extending the model from $\mathcal{A}_0$ (bounded $[0,1]$) to $\mathcal{A}$ (semi-bounded $(-\infty,1]$).

Proposed method

  • The authors refine large deviation probabilities for the empirical divergence $D_{\mathrm{inf}}(\hat{F}_i, \mu)$ to derive finite-time bounds, avoiding reliance on counting or covering empirical distributions.
  • They exploit the duality between the Kullback-Leibler divergence and the cumulant generating function via Cramér's theorem to bound tail probabilities of the empirical mean and divergence.
  • A key technical step involves bounding the probability that the empirical divergence deviates from its true value using the Legendre transform of the cumulant generating function.
  • The analysis leverages non-asymptotic bounds on the expected number of pulls of suboptimal arms by decomposing the regret into events related to overestimation and underestimation of means.
  • The method uses a novel coupling argument and moment generating function-based concentration inequalities to control the expected number of times suboptimal arms are pulled.
  • The regret bound is derived by summing over time steps and applying geometric series bounds on exponential terms arising from tail probabilities.

Experimental results

Research questions

  • RQ1Can a finite-time regret bound be established for the DMED algorithm in the semi-bounded support model where rewards are supported on $(-\infty,1]$?
  • RQ2Is the assumption of a lower-bounded support (e.g., 0 in $[0,1]$) essential for achieving the asymptotic regret bound in the DMED framework?
  • RQ3Can the regret bound of DMED be extended to distributions with only a moment generating function existing in a neighborhood of zero, rather than bounded support?
  • RQ4Does the theoretical regret lower bound remain unchanged when the model is extended from $\mathcal{A}_0$ to $\mathcal{A}$, where $\mathcal{A}_0 \subset \mathcal{A}$?

Key findings

  • The paper establishes a finite-time regret bound for the DMED algorithm in the semi-bounded support model $\mathcal{A}$, where rewards are supported on $(-\infty,1]$ and have a moment generating function in a neighborhood of zero.
  • The regret bound is of the form $\sum_{i:\mu_i < \mu^*} \frac{\log n}{D_{\mathrm{inf}}(F_i, \mu^*; \mathcal{A})} + O(1)$, matching the asymptotic lower bound up to a constant factor.
  • The authors prove that $D_{\mathrm{inf}}(F_i, \mu^*; \mathcal{A}_0) = D_{\mathrm{inf}}(F_i, \mu^*; \mathcal{A})$ for all $F_i \in \mathcal{A}_0$, showing that the theoretical lower bound is unchanged under the extended model.
  • The assumption of a lower-bounded support is not essential for the asymptotic optimality of DMED; it can be replaced by the existence of a moment generating function.
  • The finite-time regret bound is derived without counting or covering empirical distributions, avoiding the intractable terms that arise in previous non-asymptotic Sanov-based analyses.
  • The analysis shows that the expected number of pulls of suboptimal arms is bounded by a sum of exponential terms, leading to a poly-logarithmic regret bound in $n$.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.