[Paper Review] On the Optimality of Perturbations in Stochastic and Adversarial Multi-armed Bandit Problems
This paper establishes the first unified regret analysis for perturbation-based algorithms in stochastic multi-armed bandits with sub-Gaussian rewards, proving instance-optimal bounds for sub-Weibull and bounded perturbations. It further identifies fundamental barriers to achieving minimax optimality via perturbations, suggesting that optimal perturbations must be of Fréchet type, particularly in adversarial settings where Tsallis entropy regularization remains unmatched by known perturbation schemes.
We investigate the optimality of perturbation based algorithms in the stochastic and adversarial multi-armed bandit problems. For the stochastic case, we provide a unified regret analysis for both sub-Weibull and bounded perturbations when rewards are sub-Gaussian. Our bounds are instance optimal for sub-Weibull perturbations with parameter 2 that also have a matching lower tail bound, and all bounded support perturbations where there is sufficient probability mass at the extremes of the support. For the adversarial setting, we prove rigorous barriers against two natural solution approaches using tools from discrete choice theory and extreme value theory. Our results suggest that the optimal perturbation, if it exists, will be of Frechet-type.
Motivation & Objective
- To provide a unified regret analysis for perturbation-based algorithms in stochastic multi-armed bandits with sub-Gaussian rewards.
- To identify conditions under which perturbation-based algorithms achieve instance-optimal regret, particularly for sub-Weibull and bounded support perturbations.
- To investigate whether perturbations can achieve minimax optimality in adversarial bandits, matching the performance of Tsallis entropy regularization.
- To establish theoretical barriers against two natural approaches to achieving minimax-optimality via perturbations in the adversarial setting.
Proposed method
- Derives a general regret bound for Follow-The-Perturbed-Leader (FTPL) algorithms under sub-Gaussian rewards using sub-Weibull and bounded perturbations.
- Applies extreme value theory and discrete choice theory to analyze the tail behavior of perturbation distributions and their impact on regret.
- Uses the Gumbel lemma and connections between regularization and perturbation to reinterpret Thompson sampling with Gaussian priors as a perturbation algorithm.
- Employs the hazard rate condition from Abernethy et al. [2] to classify perturbation distributions and assess their suitability for adversarial bandits.
- Analyzes the choice probability function induced by Tsallis entropy and shows it cannot be exactly replicated by any perturbation in K ≥ 4 arms.
- Demonstrates that perturbations equivalent to Tsallis entropy must be of Fréchet type, based on asymptotic tail behavior and extreme value theory.
Experimental results
Research questions
- RQ1Under what conditions do perturbation-based algorithms achieve instance-optimal regret in stochastic multi-armed bandits with sub-Gaussian rewards?
- RQ2Can perturbations replicate the minimax-optimal regret of Tsallis entropy regularization in adversarial bandits?
- RQ3Are there fundamental theoretical barriers preventing perturbations from matching the performance of entropy-based regularizers like Tsallis entropy?
- RQ4What class of distributions (e.g., Fréchet, Gumbel) is necessary or sufficient for optimal perturbation-based algorithms in adversarial bandits?
- RQ5Is there a perturbation distribution that exactly emulates the choice probabilities of FTRL with Tsallis entropy regularization in multi-armed bandits?
Key findings
- The paper establishes instance-optimal regret bounds for sub-Weibull perturbations with parameter 2 and matching lower tail bounds, and for bounded support perturbations with sufficient mass at the extremes.
- For bounded support perturbations, the analysis includes the Uniform and Rademacher distributions, leading to a new regret bound for a randomized variant of UCB that selects randomly from confidence interval bounds.
- Thompson sampling with Gaussian priors and rewards is shown to be equivalent to a perturbation-based algorithm with Gaussian perturbations, generalizing prior results beyond Gaussian settings.
- The paper proves that no perturbation can exactly replicate the choice probabilities of FTRL with Tsallis entropy regularization when K ≥ 4, resolving a key open question.
- It is shown that a better analysis of existing perturbations cannot eliminate the O(√(log K)) factor in regret bounds, indicating inherent limitations in current perturbation families.
- The results suggest that the optimal perturbation, if it exists, must be of Fréchet type, based on extreme value theory and the asymptotic behavior of hazard rates in heavy-tailed distributions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.