Skip to main content
QUICK REVIEW

[Paper Review] Weak Signal Asymptotics for Sequentially Randomized Experiments

Kuang, Xu, Stefan Wager|arXiv (Cornell University)|Jan 25, 2021
Advanced Bandit Algorithms Research43 references4 citations
TL;DR

This paper introduces a weak signal asymptotic framework for sequentially randomized experiments, showing that under $1/ ext{sqrt}(n)$ scaling of reward gaps, sample paths converge weakly to a diffusion limit governed by a stochastic differential equation. The key contribution is a refined, instance-specific analysis of regret and belief dynamics, revealing that Lipschitz-continuous sampling leads to sub-optimal regret with large gaps, while Thompson sampling with asymptotically uninformative priors achieves near-optimal regret scaling even in weak signal regimes.

ABSTRACT

We use the lens of weak signal asymptotics to study a class of sequentially randomized experiments, including those that arise in solving multi-armed bandit problems. In an experiment with $n$ time steps, we let the mean reward gaps between actions scale to the order $1/\sqrt{n}$ so as to preserve the difficulty of the learning task as $n$ grows. In this regime, we show that the sample paths of a class of sequentially randomized experiments -- adapted to this scaling regime and with arm selection probabilities that vary continuously with state -- converge weakly to a diffusion limit, given as the solution to a stochastic differential equation. The diffusion limit enables us to derive refined, instance-specific characterization of stochastic dynamics, and to obtain several insights on the regret and belief evolution of a number of sequential experiments including Thompson sampling (but not UCB, which does not satisfy our continuity assumption). We show that all sequential experiments whose randomization probabilities have a Lipschitz-continuous dependence on the observed data suffer from sub-optimal regret performance when the reward gaps are relatively large. Conversely, we find that a version of Thompson sampling with an asymptotically uninformative prior variance achieves near-optimal instance-specific regret scaling, including with large reward gaps, but these good regret properties come at the cost of highly unstable posterior beliefs.

Motivation & Objective

  • To develop a refined, instance-specific understanding of sequential experiments beyond worst-case guarantees, particularly in high-stakes, small-scale applications.
  • To analyze the stochastic dynamics of adaptive experiments—especially regret and belief evolution—under a weak signal asymptotic regime where reward gaps scale as $1/ ext{sqrt}(n)$.
  • To derive a diffusion limit for sequentially randomized Markov experiments with continuous arm selection probabilities, enabling distributional insights into sample paths and performance.
  • To evaluate the performance of popular algorithms like Thompson sampling and UCB under this new asymptotic regime, particularly focusing on regret scaling and belief stability.
  • To provide theoretical justification for the use of diffuse priors in Thompson sampling, aligning with empirical folklore in bandit literature.

Proposed method

  • Adopt a weak signal asymptotic regime where the mean reward gap between arms scales as $1/ ext{sqrt}(n)$ as the number of time steps $n \to \infty$, preserving the learning difficulty.
  • Establish weak convergence of scaled sample paths of sequentially randomized Markov experiments to a diffusion process, characterized as the solution to a stochastic differential equation (SDE).
  • Use a random time change technique to express the limit cumulative reward as a time-changed Brownian motion with drift, where the time change is driven by cumulative sampling probabilities.
  • Apply the diffusion limit to analyze regret and belief evolution in Thompson sampling and Lipschitz-continuous sampling functions, focusing on instance-specific performance.
  • Characterize the impact of prior variance in Thompson sampling by comparing performance under informative vs. asymptotically uninformative priors.
  • Leverage the SDE framework to derive sharp, distributional insights into the stochastic behavior of sequential experiments, including transient belief dynamics and regret scaling.
Figure 1 : Regret profile for two-armed Thompson sampling, for $c=1$ , 1/2, 1/4, 1/8, 1/16, 1/32, 1/64, 1/256, 1/1024, and finally $c=0$ . We use $\sigma^{2}=1$ throughout. The left panel shows expected regret, while the right panel shows $\mathbb{E}\left[Q_{1}\right]$ . The curves with positive val
Figure 1 : Regret profile for two-armed Thompson sampling, for $c=1$ , 1/2, 1/4, 1/8, 1/16, 1/32, 1/64, 1/256, 1/1024, and finally $c=0$ . We use $\sigma^{2}=1$ throughout. The left panel shows expected regret, while the right panel shows $\mathbb{E}\left[Q_{1}\right]$ . The curves with positive val

Experimental results

Research questions

  • RQ1How do regret and belief dynamics in sequentially randomized experiments behave under a weak signal asymptotic regime where reward gaps scale as $1/ ext{sqrt}(n)$?
  • RQ2Under what conditions do sequentially randomized experiments converge weakly to a diffusion limit, and what is the form of the limiting SDE?
  • RQ3How does the continuity of arm selection probabilities (e.g., Lipschitz-continuous vs. discontinuous) affect regret performance in the weak signal regime?
  • RQ4Can Thompson sampling with an asymptotically uninformative prior achieve near-optimal regret scaling across all reward gap magnitudes, including large gaps?
  • RQ5To what extent do the diffusion limits derived here provide insights into human learning or scientific consensus formation in sequential information collection?

Key findings

  • Under the weak signal asymptotic regime with $1/ ext{sqrt}(n)$ scaling of reward gaps, the sample paths of sequentially randomized experiments with continuous arm selection probabilities converge weakly to a diffusion limit governed by a stochastic differential equation.
  • Lipschitz-continuous arm selection probabilities lead to non-vanishing regret even in the weak signal regime, indicating sub-optimal performance when reward gaps are relatively large.
  • Thompson sampling with an asymptotically uninformative prior achieves near-optimal instance-specific regret scaling, particularly excelling in regimes with large reward gaps.
  • The same Thompson sampling variant with uninformative priors exhibits highly unstable posterior beliefs, indicating a trade-off between regret optimality and belief stability.
  • The diffusion limit enables a refined, distributional characterization of cumulative reward and action sampling paths, offering insights beyond mean regret performance.
  • The results suggest that the use of diffuse priors in Thompson sampling is theoretically justified in weak signal settings, supporting empirical recommendations such as setting prior variance to a fixed large constant.
Figure 3 : Distribution of the (scaled) regret for two-armed Thompson sampling in the undersmoothed regime (i.e., with $c=0$ ), as a function of (scaled) arm gap $\delta$ . The histograms are aggregated over 100,000 realization of the limiting stochastic differential equation.
Figure 3 : Distribution of the (scaled) regret for two-armed Thompson sampling in the undersmoothed regime (i.e., with $c=0$ ), as a function of (scaled) arm gap $\delta$ . The histograms are aggregated over 100,000 realization of the limiting stochastic differential equation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.