[Paper Review] A Statistically Reliable Optimization Framework for Bandit Experiments in Scientific Discovery
The paper develops a general algorithm-induced test (AIT) correction to enable valid hypothesis testing under adaptive bandit sampling, and introduces an objective function to trade off reward versus statistical power, plus an optimization framework to select bandit parameters under user-specified costs.
Scientific experimentation is largely driven by statistical hypothesis testing to determine significant differences in interventions. Traditionally, experimenters allocate samples uniformly between each intervention. However, such an approach may lead to suboptimal outcomes - multi-armed bandits (MABs) addresses this problem by allocating samples adaptively to maximize outcomes. Yet, two challenges have hindered the use of MABs in scientific domains. First, common hypothesis tests (e.g., $t$-tests) become invalid under adaptive sampling without correction, leading to inflated type~I and type~II errors. This is an understudied problem, and prior solutions suffer from issues such as low statistical power which prevent adoption in many practical settings. Second, practitioners must explicitly balance cumulative reward with statistical efficiency, yet no general methodology exists to quantify this trade-off across algorithms. In this paper, we study assumption modification and critical region correction approaches for hypothesis testing that enable common tests to be applied to adaptively collected data. We provide heuristic justification for its power efficiency and show in simulation that it achieves higher power than existing approaches. Further, we derive a theoretically and practically motivated objective function for adaptive experiment evaluation, which we integrate into a unified experimental framework. Our framework asks experimenters to specify an experiment extension cost for their problem, and based on that enables our proposed optimization procedure to select the bandit algorithm that best balances reward and power in their setting. We show that our approach enables practitioners to improve outcomes with only slightly more steps than uniform randomization, while retaining statistical validity.
Motivation & Objective
- Motivate the use of adaptive (bandit) sampling to improve experimental outcomes while maintaining valid statistical inference.
- Provide a general test-correction approach that yields valid Type I error control for any bandit algorithm and common tests.
- Introduce an objective function to balance reward with statistical power under a user-defined horizon/cost.
- Develop an optimization framework to recommend bandit algorithms and experiment lengths given cost and power constraints.
- Evaluate the proposed methods through simulations across common bandit algorithms and hypothesis tests.
Proposed method
- Propose Algorithm-Induced Test (AIT) correction to construct null distributions by simulating data collection under the same adaptive algorithm and estimate a null distribution for the test statistic.
- Show that, for simple hypotheses, the LRT statistic with AIT correction is the most powerful test under adaptive data collection.
- Define and justify an experiment extension cost parameter w and derive an objective function F(T,R,w)=R/T - w*log(T) to quantify reward versus horizon.
- Formalize a PDE-based iso-value condition to justify the chosen objective and its desirable properties (monotonicity, scale/shift consistency).
- Develop an optimization procedure to select bandit algorithm parameters and horizon that maximize the proposed objective under power constraints.

Experimental results
Research questions
- RQ1How can hypothesis tests be corrected to remain valid under adaptive bandit data collection across arbitrary algorithms and tests?
- RQ2How can one quantify and optimize the trade-off between cumulative reward and statistical power in adaptive experiments?
- RQ3What algorithmic framework best balances reward and horizon given user-specified costs for extending the experiment?
- RQ4How does the proposed correction compare to existing approaches (e.g., ART) in terms of power and FPR under common bandit settings?
- RQ5Can the framework provide practical guidance for selecting bandit parameters and experiment length in real scientific settings?
Key findings
- AIT correction yields higher power than existing approaches (e.g., ART) across several algorithms (TS, ε-greedy, UCB) while maintaining empirical FPR near the target (≈0.05).
- In simple hypothesis settings, LRT with AIT correction is optimal for testing under adaptive data collection.
- The proposed ECP-reward objective F(T,R,w)=R/T - w*log(T) encodes a trade-off between average reward and experiment extension cost, with useful monotonicity and scale-shift properties.
- The framework provides an optimization toolkit that recommends bandit parameters and horizon for a given w to balance reward and statistical efficiency.
- Simulation studies indicate that the approach yields valid inference and improved practical performance with only modest increases in steps compared to uniform randomization.
- The methodology is demonstrated using common bandit algorithms (TS, ε-TS, UCB) and standard tests (t-tests, ANOVA, Tukey’s test).

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.