Skip to main content
QUICK REVIEW

[Paper Review] Optimal sequential testing of two simple hypotheses in presence of control variables

Andrey Novikov|arXiv (Cornell University)|Dec 7, 2008
Advanced Statistical Process Monitoring5 references3 citations
TL;DR

This paper develops an optimal sequential testing procedure for distinguishing between two simple hypotheses in the presence of control variables that influence the data distribution. By jointly optimizing the control policy, stopping rule, and decision rule, it establishes that the optimal strategy involves a likelihood ratio-based stopping boundary and minimizes expected sample size under Type I and II error constraints.

ABSTRACT

Suppose that at any stage of a statistical experiment a control variable $X$ that affects the distribution of the observed data $Y$ can be used. The distribution of $Y$ depends on some unknown parameter $θ$, and we consider the classical problem of testing a simple hypothesis $H_0: θ=θ_0$ against a simple alternative $H_1: θ=θ_1$ allowing the data to be controlled by $X$, in the following sequential context. The experiment starts with assigning a value $X_1$ to the control variable and observing $Y_1$ as a response. After some analysis, we choose another value $X_2$ for the control variable, and observe $Y_2$ as a response, etc. It is supposed that the experiment eventually stops, and at that moment a final decision in favour of $H_0$ or $H_1$ is to be taken. In this article, our aim is to characterize the structure of optimal sequential procedures, based on this type of data, for testing a simple hypothesis against a simple alternative.

Motivation & Objective

  • To characterize the structure of optimal sequential testing procedures when control variables influence the data distribution.
  • To minimize the expected sample size under constraints on Type I and Type II error probabilities.
  • To derive necessary and sufficient conditions for optimality of the joint control, stopping, and decision rules.
  • To extend classical sequential analysis to settings where experimental controls can be adaptively chosen to improve efficiency.
  • To establish conditions under which the optimal stopping rule has finite expectation under both hypotheses.

Proposed method

  • Models the sequential testing problem as a triplet of control policy $\chi$, stopping rule $\psi$, and decision rule $\phi$, all adapted to observed data and controls.
  • Uses a likelihood ratio process $Z_n^\chi$ to track evidence accumulation under $H_0$ and $H_1$, with optimal stopping based on thresholds $A$ and $B$.
  • Applies dynamic programming and optimal stopping theory to derive the value function and optimality conditions via recursive equations.
  • Derives the optimal stopping rule as a thresholding of the likelihood ratio, with $\psi_n^\chi$ determined by comparison with a threshold function $l_n^\chi$.
  • Introduces a Bayesian extension by minimizing a weighted sum of expected sample numbers under both hypotheses, leading to modified optimality conditions.
  • Establishes that the optimal decision rule $\phi_n$ is a threshold indicator on the likelihood ratio, rejecting $H_0$ when $\lambda_0 f_{\theta_0}^n \leq \lambda_1 f_{\theta_1}^n$.

Experimental results

Research questions

  • RQ1What is the optimal structure of a sequential testing procedure when control variables influence the data distribution?
  • RQ2How can the control policy, stopping rule, and decision rule be jointly optimized to minimize expected sample size under error constraints?
  • RQ3Under what conditions does the optimal stopping rule have finite expectation under both the null and alternative hypotheses?
  • RQ4How does the inclusion of control variables affect the likelihood ratio-based stopping boundaries compared to classical sequential analysis?
  • RQ5Can the optimality framework be extended to Bayesian settings by incorporating expected sample numbers under both hypotheses?

Key findings

  • The optimal stopping rule is characterized by a thresholding of the likelihood ratio process $Z_n^\chi$, with stopping when $Z_n^\chi$ crosses boundaries $A$ and $B$.
  • The optimal decision rule is given by $\phi_n = I_{\{\lambda_0 f_{\theta_0}^n \leq \lambda_1 f_{\theta_1}^n\}}$, which minimizes the risk under the given error constraints.
  • The optimal control policy $\chi$ ensures that the likelihood ratio process $Z_n^\chi$ is a supermartingale under $H_0$ and a submartingale under $H_1$, achieving optimality.
  • If the condition $1 + R(\infty) \geq \lambda_0$ holds, the optimal stopping rule is of the form $\tau = \inf\{n \geq 1 : Z_n^\chi \leq A\}$ or $\geq B$, ensuring finite expectation.
  • When the condition $1 + R(\infty) < \lambda_0$ holds, the optimal strategy may lead to infinite expected sample size under $H_1$, indicating potential inefficiency.
  • A Bayesian extension minimizing a weighted sum of expected sample numbers under both hypotheses leads to modified optimality conditions involving $\pi_0 E_{\theta_0}^\chi \tau_\psi + \pi_1 E_{\theta_1}^\chi \tau_\psi$.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.