[Paper Review] Weak Signal Asymptotics for Sequentially Randomized Experiments
This paper introduces a weak signal asymptotic framework for sequentially randomized experiments, showing that under $1/ ext{sqrt}(n)$ scaling of reward gaps, sample paths converge weakly to a diffusion limit governed by a stochastic differential equation. The key contribution is a refined, instance-specific analysis of regret and belief dynamics, revealing that Lipschitz-continuous sampling leads to sub-optimal regret with large gaps, while Thompson sampling with asymptotically uninformative priors achieves near-optimal regret scaling even in weak signal regimes.
We use the lens of weak signal asymptotics to study a class of sequentially randomized experiments, including those that arise in solving multi-armed bandit problems. In an experiment with $n$ time steps, we let the mean reward gaps between actions scale to the order $1/\sqrt{n}$ so as to preserve the difficulty of the learning task as $n$ grows. In this regime, we show that the sample paths of a class of sequentially randomized experiments -- adapted to this scaling regime and with arm selection probabilities that vary continuously with state -- converge weakly to a diffusion limit, given as the solution to a stochastic differential equation. The diffusion limit enables us to derive refined, instance-specific characterization of stochastic dynamics, and to obtain several insights on the regret and belief evolution of a number of sequential experiments including Thompson sampling (but not UCB, which does not satisfy our continuity assumption). We show that all sequential experiments whose randomization probabilities have a Lipschitz-continuous dependence on the observed data suffer from sub-optimal regret performance when the reward gaps are relatively large. Conversely, we find that a version of Thompson sampling with an asymptotically uninformative prior variance achieves near-optimal instance-specific regret scaling, including with large reward gaps, but these good regret properties come at the cost of highly unstable posterior beliefs.
Motivation & Objective
- To develop a refined, instance-specific understanding of sequential experiments beyond worst-case guarantees, particularly in high-stakes, small-scale applications.
- To analyze the stochastic dynamics of adaptive experiments—especially regret and belief evolution—under a weak signal asymptotic regime where reward gaps scale as $1/ ext{sqrt}(n)$.
- To derive a diffusion limit for sequentially randomized Markov experiments with continuous arm selection probabilities, enabling distributional insights into sample paths and performance.
- To evaluate the performance of popular algorithms like Thompson sampling and UCB under this new asymptotic regime, particularly focusing on regret scaling and belief stability.
- To provide theoretical justification for the use of diffuse priors in Thompson sampling, aligning with empirical folklore in bandit literature.
Proposed method
- Adopt a weak signal asymptotic regime where the mean reward gap between arms scales as $1/ ext{sqrt}(n)$ as the number of time steps $n \to \infty$, preserving the learning difficulty.
- Establish weak convergence of scaled sample paths of sequentially randomized Markov experiments to a diffusion process, characterized as the solution to a stochastic differential equation (SDE).
- Use a random time change technique to express the limit cumulative reward as a time-changed Brownian motion with drift, where the time change is driven by cumulative sampling probabilities.
- Apply the diffusion limit to analyze regret and belief evolution in Thompson sampling and Lipschitz-continuous sampling functions, focusing on instance-specific performance.
- Characterize the impact of prior variance in Thompson sampling by comparing performance under informative vs. asymptotically uninformative priors.
- Leverage the SDE framework to derive sharp, distributional insights into the stochastic behavior of sequential experiments, including transient belief dynamics and regret scaling.
![Figure 1 : Regret profile for two-armed Thompson sampling, for $c=1$ , 1/2, 1/4, 1/8, 1/16, 1/32, 1/64, 1/256, 1/1024, and finally $c=0$ . We use $\sigma^{2}=1$ throughout. The left panel shows expected regret, while the right panel shows $\mathbb{E}\left[Q_{1}\right]$ . The curves with positive val](https://ar5iv.labs.arxiv.org/html/2101.09855/assets/x1.png)
Experimental results
Research questions
- RQ1How do regret and belief dynamics in sequentially randomized experiments behave under a weak signal asymptotic regime where reward gaps scale as $1/ ext{sqrt}(n)$?
- RQ2Under what conditions do sequentially randomized experiments converge weakly to a diffusion limit, and what is the form of the limiting SDE?
- RQ3How does the continuity of arm selection probabilities (e.g., Lipschitz-continuous vs. discontinuous) affect regret performance in the weak signal regime?
- RQ4Can Thompson sampling with an asymptotically uninformative prior achieve near-optimal regret scaling across all reward gap magnitudes, including large gaps?
- RQ5To what extent do the diffusion limits derived here provide insights into human learning or scientific consensus formation in sequential information collection?
Key findings
- Under the weak signal asymptotic regime with $1/ ext{sqrt}(n)$ scaling of reward gaps, the sample paths of sequentially randomized experiments with continuous arm selection probabilities converge weakly to a diffusion limit governed by a stochastic differential equation.
- Lipschitz-continuous arm selection probabilities lead to non-vanishing regret even in the weak signal regime, indicating sub-optimal performance when reward gaps are relatively large.
- Thompson sampling with an asymptotically uninformative prior achieves near-optimal instance-specific regret scaling, particularly excelling in regimes with large reward gaps.
- The same Thompson sampling variant with uninformative priors exhibits highly unstable posterior beliefs, indicating a trade-off between regret optimality and belief stability.
- The diffusion limit enables a refined, distributional characterization of cumulative reward and action sampling paths, offering insights beyond mean regret performance.
- The results suggest that the use of diffuse priors in Thompson sampling is theoretically justified in weak signal settings, supporting empirical recommendations such as setting prior variance to a fixed large constant.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.