[Paper Review] Non-Asymptotic Sequential Tests for Overlapping Hypotheses and application to near optimal arm identification in bandit models
This paper develops non-asymptotic sequential tests for overlapping hypotheses using a parallel Generalized Likelihood Ratio Test (GLRT), applying it to PAC-best arm identification in multi-armed bandits. It establishes tight non-asymptotic bounds on sample complexity and provides a lower bound based on information-theoretic arguments, showing that the proposed strategy achieves optimal performance for regular bandit instances under (ϵ,δ)-PAC criteria.
In this paper, we study sequential testing problems with \emph{overlapping} hypotheses. We first focus on the simple problem of assessing if the mean $\mu$ of a Gaussian distribution is smaller or larger than a fixed $\epsilon>0$; if $\mu\in(-\epsilon,\epsilon)$, both answers are considered to be correct. Then, we consider PAC-best arm identification in a bandit model: given $K$ probability distributions on $\mathbb{R}$ with means $\mu_1,\dots,\mu_K$, we derive the asymptotic complexity of identifying, with risk at most $\delta$, an index $I\in\{1,\dots,K\}$ such that $\mu_I\geq \max_i\mu_i -\epsilon$. We provide non-asymptotic bounds on the error of a parallel General Likelihood Ratio Test, which can also be used for more general testing problems. We further propose lower bound on the number of observation needed to identify a correct hypothesis. Those lower bounds rely on information-theoretic arguments, and specifically on two versions of a change of measure lemma (a high-level form, and a low-level form) whose relative merits are discussed.
Motivation & Objective
- To address the lack of non-asymptotic analysis for sequential testing under overlapping hypotheses, especially in active learning settings.
- To generalize δ-correct best-arm identification to the more practical (ϵ,δ)-PAC framework, where multiple arms may be nearly optimal.
- To derive non-asymptotic lower bounds on sample complexity using information-theoretic tools, particularly change-of-measure lemmas.
- To design and analyze a (ϵ,δ)-PAC strategy combining parallel GLRT with a tracking sampling rule that matches the lower bound for regular bandit models.
Proposed method
- Proposes a parallel Generalized Likelihood Ratio Test (GLRT) for sequential testing under overlapping hypotheses, where multiple hypotheses may be simultaneously valid.
- Uses a change-of-measure argument with two forms—high-level and low-level—to derive non-asymptotic error bounds for the GLRT.
- Introduces a tracking sampling rule that adaptively allocates pulls to arms to converge toward an optimal allocation for minimizing sample complexity.
- Derives a non-asymptotic lower bound on the expected stopping time using a convex optimization formulation of the characteristic time.
- Applies the GLRT to multi-armed bandit models with Gaussian or exponential family distributions, ensuring δ-correctness under (ϵ,δ)-PAC criteria.
- Solves for optimal sampling weights via a min-max optimization problem involving Kullback-Leibler divergences, with a closed-form solution under certain conditions.
Experimental results
Research questions
- RQ1Can non-asymptotic bounds be derived for sequential testing when hypotheses overlap, i.e., when multiple hypotheses may be correct?
- RQ2What is the minimal sample complexity required to identify an ϵ-optimal arm in a multi-armed bandit model with high probability (1−δ)?
- RQ3How does the performance of the parallel GLRT compare to the information-theoretic lower bound in the (ϵ,δ)-PAC setting?
- RQ4Can a tracking sampling rule be designed to achieve the lower bound on sample complexity for regular bandit instances?
- RQ5What is the structure of the optimal sampling weights in the (ϵ,δ)-PAC setting, and how can they be computed efficiently?
Key findings
- The parallel GLRT achieves δ-correctness for (ϵ,δ)-PAC best-arm identification in multi-armed bandits, with non-asymptotic error bounds that scale as ln(1/δ).
- A non-asymptotic lower bound on the expected stopping time is derived using information-theoretic arguments, showing that the characteristic time is the solution to a convex optimization problem.
- For regular bandit instances, the proposed (ϵ,δ)-PAC strategy with tracking sampling rule matches the derived lower bound, proving asymptotic optimality.
- The optimal sampling weights are computed as the solution to a min-max optimization problem involving KL divergences, with a unique solution characterized by a system of equations involving inverse KL derivatives.
- In the special case where an arm has mean µ+ −ϵ and another has mean µ+, the optimal weight assigns full allocation to the best arm, simplifying the solution.
- The analysis reveals that high-level information-theoretic reasoning fails in the overlapping hypothesis setting, necessitating low-level change-of-measure arguments for tight bounds.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.