Skip to main content
QUICK REVIEW

[Paper Review] Asymptotic Representations for Sequential Decisions, Adaptive Experiments, and Batched Bandits

Keisuke Hirano, Jack R. Porter|arXiv (Cornell University)|Feb 6, 2023
Advanced Statistical Process MonitoringDecision Sciences3 citations
TL;DR

This paper develops asymptotic representation theory for sequential decision problems involving adaptive experiments and batched bandits, showing that all limit distributions under local alternatives can be characterized by a single limiting Gaussian bandit. The framework enables valid inference, optimal design search, and power analysis in adaptive settings where treatment allocation affects later data collection.

ABSTRACT

We develop asymptotic approximations that can be applied to sequential estimation and inference problems, adaptive randomized controlled trials, and related settings. In batched adaptive settings where the decision at one stage can affect the observation of variables in later stages, our asymptotic representation characterizes all limit distributions attainable through a joint choice of an adaptive design rule and statistics applied to the adaptively generated data. This facilitates local power analysis of tests, comparison of adaptive treatments rules, and other analyses of batchwise sequential statistical decision rules.

Motivation & Objective

  • To address the lack of general asymptotic approximation tools for sequential and adaptive statistical decision problems with structured, endogenous information sets.
  • To extend classic asymptotic representation theorems to dynamic, multi-stage settings where treatment allocation and data collection are adaptively chosen.
  • To provide a unified framework for analyzing the asymptotic behavior of statistical procedures in adaptive randomized controlled trials and batched bandit problems.
  • To enable the comparison of adaptive sampling rules and inference procedures under local asymptotic theory.
  • To support the design of optimal adaptive experiments by characterizing attainable limit distributions through a limiting data environment (limit bandit).

Proposed method

  • Develop a limiting data environment—called a 'limit bandit'—that captures sequential informational constraints and interactions between adaptive design rules and data generation.
  • Use local parametrization in the spirit of Hirano and Porter (2009) to model local alternatives where treatment effects are small but estimable.
  • Construct asymptotic representations that link the original sequential decision problem to a simplified limiting Gaussian bandit model.
  • Apply the representation to analyze the local asymptotic power of tests such as the ZJM test and pooled difference-in-means under Thompson sampling.
  • Use numerical evaluation (simulation or integration) to compare power curves of different inference procedures under the same limiting model.
  • Introduce a weighted Thompson sampling rule that shrinks toward balanced allocation to assess trade-offs between allocative efficiency and inferential efficiency.

Experimental results

Research questions

  • RQ1How can asymptotic approximations be generalized to sequential decision problems with adaptive sampling and treatment allocation?
  • RQ2What is the limiting distribution of test statistics in adaptive experiments where treatment assignment depends on past data?
  • RQ3How does the power of inference procedures like the ZJM test compare to optimal power envelopes in adaptive batched designs?
  • RQ4To what extent does shrinking Thompson sampling toward balanced allocation improve statistical power for treatment effect testing?
  • RQ5Can the limiting Gaussian bandit representation be used to compare and optimize adaptive design rules and inference procedures?

Key findings

  • The limiting Gaussian bandit representation characterizes all attainable limit distributions in sequential adaptive experiments, enabling a unified framework for inference and design.
  • The ZJM test, which uses only batchwise differences in means, achieves power close to the limited power envelope for two-dimensional statistics under small alternatives.
  • The pooled difference-in-means test exhibits uniformly higher power than the ZJM test and approaches the full four-dimensional data power envelope.
  • Shrinking the Thompson sampling rule toward 50-50 allocation (e.g., with c=0.1) reduces the power gap between the ZJM test and the optimal power envelope, though the relative ordering of test performance is preserved.
  • The limiting bandit framework allows computationally tractable comparison of inference procedures and adaptive rules without simulating the full original data-generating process.
  • The representation is general and does not require restricting to specific rules (e.g., Bayes or plug-in), enabling application to minmax regret or robust preference models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.