[Paper Review] Generic E-Variables for Exact Sequential k-Sample Tests that allow for Optional Stopping
This paper introduces generic E-variables for exact, nonasymptotic k-sample sequential tests that maintain type-I error control under optional stopping and continuation. By constructing block-wise E-variables and multiplying them into a test martingale, the method ensures valid inference across flexible sampling designs, with optimal growth properties under the alternative when using relative GRO priors, outperforming p-value methods in early stopping scenarios while preserving error guarantees.
We develop E-variables for testing whether two or more data streams come from the same source or not, and more generally, whether the difference between the sources is larger than some minimal effect size. These E-variables lead to exact, nonasymptotic tests that remain safe, i.e. keep their type-I error guarantees, under flexible sampling scenarios such as optional stopping and continuation. In special cases our E-variables also have an optimal 'growth' property under the alternative. While the construction is generic, we illustrate it through the special case of k x 2 contingency tables, where we also allow for the incorporation of different restrictions on a composite alternative. Comparison to p-value analysis in simulations and a real-world example show that E-variables, through their flexibility, often allow for early stopping of data collection, thereby retaining similar power as classical methods, while also retaining the option of extending or combining data afterwards.
Motivation & Objective
- To develop exact, nonasymptotic hypothesis tests for k-sample comparisons that remain valid under flexible sampling, including optional stopping and continuation.
- To ensure type-I error control regardless of stopping rule, even when data collection is adapted based on observed results.
- To provide a generic construction of E-variables applicable to any exponential family and divergence measure, enabling broad applicability.
- To optimize power under the alternative via growth-rate optimal (GRO) priors, particularly in the k×2 contingency table setting.
- To demonstrate through simulations and real-world examples that E-variables allow earlier stopping with comparable power to classical p-value methods.
Proposed method
- Construct E-variables for a single block of data from two or more groups, ensuring E[s] ≤ 1 under the null hypothesis.
- Extend block E-variables to sequential data by multiplying them into a test martingale, which remains a valid E-variable under optional stopping.
- Use a prior distribution on the alternative hypothesis space to define the E-variable, with choice affecting power but not type-I error control.
- Apply the relative GRO (relative growth-rate optimality) criterion to select priors that maximize expected growth under the alternative.
- For k×2 contingency tables, derive E-variables based on conjugate priors and block-wise sampling, allowing for restrictions on effect size δ.
- Use the E-value as a test statistic: reject the null if the product exceeds 1/α, ensuring level-α error control.
Experimental results
Research questions
- RQ1Can E-variables be constructed generically for k-sample sequential tests that maintain type-I error control under optional stopping?
- RQ2How does the choice of prior on the alternative affect the growth rate and power of E-variables in sequential testing?
- RQ3In what ways do E-variables outperform p-values in terms of early stopping and power when optional stopping is applied?
- RQ4Can E-variables be constructed to have optimal growth under the alternative, particularly in the k×2 contingency table setting?
- RQ5Do E-variables based on Gunel-Dickey Bayes factors qualify as valid E-variables under the null, and if not, why?
Key findings
- E-variables constructed via block-wise multiplication maintain type-I error control under optional stopping, even when data collection is adapted based on past results.
- Simulations show that E-variables allow earlier stopping with power comparable to Fisher’s exact test, while avoiding the inflated type-I error seen in p-value methods under optional stopping.
- The E-variable based on the independent multinomial sampling scheme with a uniform prior (β = 1/2) maintains E[s] ≤ 1 under the null, satisfying the E-variable condition.
- The Gunel-Dickey Bayes factor for the Poisson sampling scheme does not qualify as an E-variable, as its expectation under the null exceeds 1 for λ ≥ 1.
- For the independent multinomial case, numerical simulations confirm that the Bayes factor's expectation exceeds 1 under the null for various θ and n_g, disqualifying it as an E-variable.
- The type-I error rate of E-variables remains bounded at α = 0.05 under optional stopping, while Fisher’s p-value method shows a sharp increase in error rate over time.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.