[Paper Review] Optimal Best Arm Identification with Fixed Confidence
The paper derives a tight non-asymptotic lower bound for best arm identification in one-parameter bandits and introduces the Track-and-Stop strategy, proven asymptotically optimal for fixed-confidence settings.
We give a complete characterization of the complexity of best-arm identification in one-parameter bandit problems. We prove a new, tight lower bound on the sample complexity. We propose the `Track-and-Stop' strategy, which we prove to be asymptotically optimal. It consists in a new sampling rule (which tracks the optimal proportions of arm draws highlighted by the lower bound) and in a stopping rule named after Chernoff, for which we give a new analysis.
Motivation & Objective
- Characterize the exact sample complexity required for delta-PAC best arm identification in one-parameter exponential families.
- Provide a tight non-asymptotic lower bound on the expected number of samples.
- Propose a learning strategy (Track-and-Stop) that achieves the lower bound asymptotically.
- Analyze stopping rules and sampling schemes that ensure delta-PAC guarantees.
Proposed method
- Derive a tight lower bound involving a problem-specific characteristic time T*(mu) through a transportation-based change of measure.
- Define the optimal arm sampling proportions w*(mu) by solving an optimization over the alternative models Alt(mu).
- Introduce the Track-and-Stop algorithm consisting of a sampling rule that tracks the optimal proportions and a Chernoff-type stopping rule with a tunable threshold.
- Provide two tracking schemes (C-Tracking and D-Tracking) that enforce exploration to guarantee convergence of empirical means.
- Analyze the stopping rule through generalized likelihood ratio statistics Z_{a,b}(t) and show how the threshold beta(t, delta) yields delta-PAC guarantees.
- Offer MDL interpretations and connect stopping behavior to information-theoretic coding arguments.
Experimental results
Research questions
- RQ1What is the correct problem-dependent lower bound on the expected sample complexity for delta-PAC best arm identification in exponential-family bandits?
- RQ2How can one compute the optimal arm-sampling proportions w*(mu) and the corresponding characteristic time T*(mu)?
- RQ3Can a practical strategy (Track-and-Stop) attain the lower bound asymptotically while satisfying delta-PAC constraints?
- RQ4How should stopping and sampling rules be designed to ensure fixed-confidence guarantees across a broad class of bandit models?
- RQ5What interpretations (statistical, information-theoretic, MDL) illuminate the stopping rule?
Key findings
- A tight, non-asymptotic lower bound on E_mu[tau_delta] is established, involving a problem-dependent characteristic time T*(mu).
- An explicit characterization of the optimal sampling proportions w*(mu) is given, enabling tracking-based strategies to achieve the lower bound.
- The Track-and-Stop strategy is proposed and shown to be asymptotically optimal as delta -> 0 under delta-PAC constraints.
- Two practical tracking schemes (C-Tracking and D-Tracking) are proven to ensure convergence of empirical means to the optimal proportions and satisfy delta-PAC.
- A Chernoff-type stopping rule with an MDL/information-theoretic interpretation yields a stopping time that achieves the lower bound in expectation up to log(1/delta) factors.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.