[Paper Review] Structural Properties of Bayesian Bandits with Exponential Family Distributions
This paper establishes general structural properties of Bayesian bandits with exponential family likelihoods and conjugate priors, proving that an arm's value increases with its prior mean and decreases with its prior weight—unifying and extending known results for Bernoulli and normal bandits. The key contribution is a formal, distribution-free characterization of the exploration-exploitation trade-off via stochastic ordering and convexity.
We study a bandit problem where observations from each arm have an exponential family distribution and different arms are assigned independent conjugate priors. At each of n stages, one arm is to be selected based on past observations. The goal is to find a strategy that maximizes the expected discounted sum of the $n$ observations. Two structural results hold in broad generality: (i) for a fixed prior weight, an arm becomes more desirable as its prior mean increases; (ii) for a fixed prior mean, an arm becomes more desirable as its prior weight decreases. These generalize and unify several results in the literature concerning specific problems including Bernoulli and normal bandits. The second result captures an aspect of the exploration-exploitation dilemma in precise terms: given the same immediate payoff, the less one knows about an arm, the more desirable it becomes because there remains more information to be gained when selecting that arm. For Bernoulli and normal bandits we also obtain extensions to nonconjugate priors.
Motivation & Objective
- To unify and generalize monotonicity results for Bayesian bandits across exponential family distributions.
- To formalize the exploration-exploitation trade-off in terms of prior mean and prior weight using stochastic ordering.
- To extend structural results beyond conjugate priors for Bernoulli and normal bandits.
- To provide a theoretical foundation for the Gittins index and dynamic allocation in finite-horizon, non-geometric discounting settings.
Proposed method
- Formalizing the two-armed bandit problem using exponential family distributions with conjugate priors parameterized by prior mean γ and prior weight τ.
- Using convex order and likelihood ratio order to compare posterior distributions and value functions across different priors.
- Applying backward induction and dynamic programming to derive value functions for finite-horizon problems with general discount sequences.
- Proving that the value function is increasing in prior mean and decreasing in prior weight via induction and stochastic dominance arguments.
- Extending results to nonconjugate priors in Bernoulli and normal bandits using convex order and posterior variance bounds.
- Leveraging properties of normal location-scale families and posterior mean derivatives to establish monotonicity under convolution.
Experimental results
Research questions
- RQ1How does the prior mean affect the optimal value of an arm in a Bayesian bandit with exponential family likelihoods?
- RQ2How does the prior weight (sample size) influence the desirability of an arm in the presence of uncertainty?
- RQ3Can the exploration-exploitation trade-off be characterized uniformly across exponential family bandits using stochastic ordering?
- RQ4Does the Gittins index monotonicity extend to non-geometric discounting and general conjugate priors?
- RQ5Can monotonicity results be extended to nonconjugate priors in Bernoulli and normal bandits?
Key findings
- For fixed prior weight, the optimal expected payoff increases as the prior mean of any arm increases.
- For fixed prior mean, the optimal expected payoff increases as the prior weight of any arm decreases, formalizing the value of exploration.
- The value function is convex in the mean of a normal prior, supporting the use of posterior mean as a sufficient statistic.
- The posterior mean derivative equals the posterior variance, enabling bounds on how quickly beliefs update.
- For nonconjugate priors in Bernoulli and normal bandits, the same monotonicity holds under convex order and stochastic dominance.
- The conjecture that a lower prior weight always yields a higher Gittins index remains open for non-geometric discounting.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.