[Paper Review] Some Notes on Blinded Sample Size Re-Estimation
This paper demonstrates that blinded sample size re-estimation (SSR) methods, particularly the Kieser and Friede approach, can lead to type I error inflation in small-sample superiority trials and significantly higher inflation in non-inferiority testing. The authors show that the method's reliance on variance estimates under the null hypothesis fails to control the type I error rate in certain borderline cases, especially when sample sizes are small, and explain why this issue was previously overlooked in the literature.
This note investigates a number of scenarios in which unadjusted testing following a blinded sample size re-estimation leads to type I error violations. For superiority testing, this occurs in certain small-sample borderline cases. We discuss a number of alternative approaches that keep the type I error rate. The paper also gives a reason why the type I error inflation in the superiority context might have been missed in previous publications and investigates why it is more marked in case of non-inferiority testing.
Motivation & Objective
- To investigate scenarios where blinded sample size re-estimation leads to type I error rate violations in clinical trials.
- To explain why type I error inflation in superiority testing is small but non-negligible in small-sample borderline cases.
- To analyze why type I error inflation is more severe in non-inferiority testing compared to superiority testing.
- To clarify why previous studies may have missed this issue, particularly due to reliance on simulation-based evidence without formal proof.
- To evaluate the implications of using blinded SSR in confirmatory clinical trials, especially in non-inferiority settings.
Proposed method
- Uses a one-sample t-test framework to model blinded SSR, focusing on sample size adjustment after n₁ = 2 observations based on total variance.
- Applies geometric probability arguments in three-dimensional space to compute rejection regions and type I error rates under H₀.
- Compares the rejection region of the full procedure (n=2 or n=3) to the union of spherical cone (n=3) and cylindrical segment (n=2), subtracting their intersection.
- Derives the distribution of the test statistic under H₀ using non-central chi-square approximations for the sum of squares.
- Analyzes the non-inferiority setting by transforming the original data via a shift δ, showing equivalence to unblinded superiority testing.
- Demonstrates that the sample size re-estimation rule in non-inferiority testing depends on the observed mean difference, introducing bias even when blinded.
Experimental results
Research questions
- RQ1Under what conditions does blinded sample size re-estimation using the Kieser and Friede method lead to type I error rate inflation in superiority trials?
- RQ2Why was the type I error inflation in superiority testing previously overlooked in the literature despite simulation evidence?
- RQ3Why is type I error inflation more pronounced in non-inferiority testing compared to superiority testing?
- RQ4How does the dependence of sample size re-estimation on the observed mean difference in non-inferiority trials affect type I error control?
- RQ5Can alternative methods such as data-dependent weighting of stage statistics provide valid type I error control in blinded SSR?
Key findings
- Blinded sample size re-estimation using the Kieser and Friede method leads to type I error inflation in small-sample superiority trials, particularly when n₁ = 2 and the variance estimate is low.
- The type I error inflation is extremely small and asymptotically vanishes as sample sizes increase, but remains non-zero in finite samples.
- The reason for the inflation lies in the geometric overlap between the rejection regions of the n=2 and n=3 cases, which causes a non-uniform distribution of type I error under H₀.
- In non-inferiority testing, the method is equivalent to an unblinded superiority test because the sample size re-estimation rule depends on the observed mean difference, leading to substantial type I error inflation.
- The distribution of the test statistic under H₀ is not standard t when sample size is data-dependent, invalidating the use of standard critical values.
- While asymptotic type I error control holds, finite-sample type I error violations are non-negligible in practical settings, especially in non-inferiority trials, making blinded SSR unsuitable for confirmatory trials.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.