Skip to main content
QUICK REVIEW

[论文解读] Sequential Nonparametric Testing with the Law of the Iterated Logarithm

Akshay Balsubramani, Aaditya Ramdas|arXiv (Cornell University)|Jun 10, 2015
Statistical Methods in Clinical Trials参考文献 22被引用 11
一句话总结

本文提出了一种新颖的独立同分布数据序列非参数检验框架,实现线性时间、常数空间计算,并具备有限样本的 I 类与 II 类错误控制。通过构建零均值鞅检验统计量,并利用拒绝阈值的统一非渐近迭代对数律(LIL),该方法能自适应地在最少样本后停止——其功效与批量检验相当,且无需先验知识即可自动检测问题的难度。

ABSTRACT

We propose a new algorithmic framework for sequential hypothesis testing with i.i.d. data, which includes A/B testing, nonparametric two-sample testing, and independence testing as special cases. It is novel in several ways: (a) it takes linear time and constant space to compute on the fly, (b) it has the same power guarantee as a non-sequential version of the test with the same computational constraints up to a small factor, and (c) it accesses only as many samples as are required - its stopping time adapts to the unknown difficulty of the problem. All our test statistics are constructed to be zero-mean martingales under the null hypothesis, and the rejection threshold is governed by a uniform non-asymptotic law of the iterated logarithm (LIL). For the case of nonparametric two-sample mean testing, we also provide a finite sample power analysis, and the first non-asymptotic stopping time calculations for this class of problems. We verify our predictions for type I and II errors and stopping times using simulations.

研究动机与目标

  • 为复合原假设下的多变量非参数问题开发一种序列假设检验框架,计算与存储开销最小化。
  • 在不依赖中心极限定理(CLT)等渐近近似的情况下,确保有限样本的 I 类错误控制。
  • 通过自适应调整停止时间以匹配问题的未知难度,实现接近最优的 II 类错误控制。
  • 实现理论保证下的早期停止,其功效与线性时间批量检验相当,同时仅使用必要数量的样本。
  • 提供一种通用且可扩展的方法,适用于独立同分布数据下的 A/B 测试、两样本检验与独立性检验。

提出的方法

  • 该方法在原假设下构建零均值鞅检验统计量,确保序列推断的有效性。
  • 利用统一的非渐近迭代对数律(LIL)版本,设定拒绝阈值,实现对时间上统一的 I 类错误控制。
  • 停止时间由检验统计量首次超出 LIL 区间的时间决定,该区间自适应于实际信号强度。
  • 该框架利用二阶统一鞅集中不等式(如 Freedman 不等式)推导出紧致的有限样本偏差界。
  • 通过在经验方差上按指数增长的周期进行剥皮论证,实现浓度中的最优迭代对数率。
  • 该方法仅依赖于在线计算的标量检验统计量,每条观测仅需常数空间与线性时间。

实验结果

研究问题

  • RQ1序列非参数检验能否在不依赖渐近正态近似的情况下,实现有限样本的 I 类错误控制?
  • RQ2如何自适应地控制停止时间,以匹配未知问题难度,而无需先验知识?
  • RQ3与批量检验相比,序列检验的有限样本功效行为如何?
  • RQ4迭代对数律能否用于推导高维或非参数设定下序列检验的非渐近统一界?
  • RQ5序列检验的停止时间在多大程度上反映了真实信号强度?其与最优批量采样相比如何?

主要发现

  • 所提方法利用有限样本 LIL 界,在所有停止时间上统一控制 I 类错误,避免依赖渐近正态性。
  • 该检验与线性时间批量检验相比,功效接近最优,且功效损失被限制在一个较小常数因子内。
  • 停止时间自适应于未知问题难度,仅当所需 II 类错误趋近于零(即功效趋近于 1)时才发散。
  • 对于非参数两样本均值检验,本文首次提供了有限样本功效分析与非渐近停止时间计算。
  • 该方法实现了最优浓度率 $ Oig(ig(\text{Var}_n \big)\text{ln ln} \frac{\text{Var}_n}{\theta}\big)^{1/2} $,与 LIL 的理论极限一致。
  • 模拟结果证实,预测的 I 类与 II 类错误率及停止时间与实际性能高度吻合。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。