Skip to main content
QUICK REVIEW

[Paper Review] Optimal Best Arm Identification with Fixed Confidence

Aurélien Garivier, Emilie Kaufmann|arXiv (Cornell University)|Feb 15, 2016
Advanced Bandit Algorithms Research40 references101 citations
TL;DR

The paper derives a tight non-asymptotic lower bound for best arm identification in one-parameter bandits and introduces the Track-and-Stop strategy, proven asymptotically optimal for fixed-confidence settings.

ABSTRACT

We give a complete characterization of the complexity of best-arm identification in one-parameter bandit problems. We prove a new, tight lower bound on the sample complexity. We propose the `Track-and-Stop' strategy, which we prove to be asymptotically optimal. It consists in a new sampling rule (which tracks the optimal proportions of arm draws highlighted by the lower bound) and in a stopping rule named after Chernoff, for which we give a new analysis.

Motivation & Objective

  • Characterize the exact sample complexity required for delta-PAC best arm identification in one-parameter exponential families.
  • Provide a tight non-asymptotic lower bound on the expected number of samples.
  • Propose a learning strategy (Track-and-Stop) that achieves the lower bound asymptotically.
  • Analyze stopping rules and sampling schemes that ensure delta-PAC guarantees.

Proposed method

  • Derive a tight lower bound involving a problem-specific characteristic time T*(mu) through a transportation-based change of measure.
  • Define the optimal arm sampling proportions w*(mu) by solving an optimization over the alternative models Alt(mu).
  • Introduce the Track-and-Stop algorithm consisting of a sampling rule that tracks the optimal proportions and a Chernoff-type stopping rule with a tunable threshold.
  • Provide two tracking schemes (C-Tracking and D-Tracking) that enforce exploration to guarantee convergence of empirical means.
  • Analyze the stopping rule through generalized likelihood ratio statistics Z_{a,b}(t) and show how the threshold beta(t, delta) yields delta-PAC guarantees.
  • Offer MDL interpretations and connect stopping behavior to information-theoretic coding arguments.

Experimental results

Research questions

  • RQ1What is the correct problem-dependent lower bound on the expected sample complexity for delta-PAC best arm identification in exponential-family bandits?
  • RQ2How can one compute the optimal arm-sampling proportions w*(mu) and the corresponding characteristic time T*(mu)?
  • RQ3Can a practical strategy (Track-and-Stop) attain the lower bound asymptotically while satisfying delta-PAC constraints?
  • RQ4How should stopping and sampling rules be designed to ensure fixed-confidence guarantees across a broad class of bandit models?
  • RQ5What interpretations (statistical, information-theoretic, MDL) illuminate the stopping rule?

Key findings

  • A tight, non-asymptotic lower bound on E_mu[tau_delta] is established, involving a problem-dependent characteristic time T*(mu).
  • An explicit characterization of the optimal sampling proportions w*(mu) is given, enabling tracking-based strategies to achieve the lower bound.
  • The Track-and-Stop strategy is proposed and shown to be asymptotically optimal as delta -> 0 under delta-PAC constraints.
  • Two practical tracking schemes (C-Tracking and D-Tracking) are proven to ensure convergence of empirical means to the optimal proportions and satisfy delta-PAC.
  • A Chernoff-type stopping rule with an MDL/information-theoretic interpretation yields a stopping time that achieves the lower bound in expectation up to log(1/delta) factors.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.