[论文解读] Time-uniform central limit theory and asymptotic confidence sequences
本文提出了渐近置信序列(AsympCSs),这是一种新颖的时间一致框架,将经典的中心极限定理(CLT)推断扩展至序列设置。通过利用Strassen和Komlós-Major-Tusnády的强不变性原理,该方法在与CLT相同的弱矩条件下,构建了普遍有效的置信序列,从而在不牺牲渐近效率的前提下,实现了对因果推断中观测研究和随机化研究的连续监测与自适应停止。
Confidence intervals based on the central limit theorem (CLT) are a cornerstone of classical statistics. Despite being only asymptotically valid, they are ubiquitous because they permit statistical inference under weak assumptions and can often be applied to problems even when nonasymptotic inference is impossible. This paper introduces time-uniform analogues of such asymptotic confidence intervals, adding to the literature on confidence sequences (CS) -- sequences of confidence intervals that are uniformly valid over time -- which provide valid inference at arbitrary stopping times and incur no penalties for "peeking" at the data, unlike classical confidence intervals which require the sample size to be fixed in advance. Existing CSs in the literature are nonasymptotic, enjoying finite-sample guarantees but not the aforementioned broad applicability of asymptotic confidence intervals. This work provides a definition for "asymptotic CSs" and a general recipe for deriving them. Asymptotic CSs forgo nonasymptotic validity for CLT-like versatility and (asymptotic) time-uniform guarantees. While the CLT approximates the distribution of a sample average by that of a Gaussian for a fixed sample size, we use strong invariance principles (stemming from the seminal 1960s work of Strassen) to uniformly approximate the entire sample average process by an implicit Gaussian process. As an illustration, we derive asymptotic CSs for the average treatment effect in observational studies (for which nonasymptotic bounds are essentially impossible to derive even in the fixed-time regime) as well as randomized experiments, enabling causal inference in sequential environments.
研究动机与目标
- 通过定义渐近置信序列(AsympCSs),弥合渐近置信区间与时间一致推断之间的差距。
- 开发一种普遍适用的AsympCS,使其在与经典中心极限定理相同的弱矩假设下保持有效性。
- 在因果推断问题中实现连续、数据自适应的监测与推断——尤其在非渐近界不可得的观测研究中。
- 提供一种框架,结合渐近推断的广泛适用性与置信序列的灵活性,避免因频繁查看或提前停止而产生惩罚。
提出的方法
- 利用强不变性原理(Strassen,Komlós-Major-Tusnády)在时间上统一近似样本均值过程为高斯过程。
- 应用Robbins的正态混合边界,构建具有精确第一类错误控制的时间一致置信序列。
- 采用序列样本分割和交叉拟合方法,在半参数模型中估计干扰函数,同时保持渐近有效性。
- 通过高效影响函数推导平均处理效应的渐近置信序列,使用 $σ^2_t = \widehat{\mathrm{var}}_t(\widehat{f})$ 进行方差估计。
- 证明AsympCS的宽度与高效影响函数的估计标准差成比例,从而在非渐近设置下也实现方差自适应。
- 证明所得到的AsympCS在关键组成部分上不可改进:正态混合边界、近似误差率和标准误估计,这是由强不变性原理带来的根本限制。
实验结果
研究问题
- RQ1能否构造一种时间一致的置信序列,使其在与经典中心极限定理相同的弱矩条件下保持渐近有效性?
- RQ2是否可能在不依赖非渐近界的前提下,构造一种在数据依赖停止时间下仍保持有效性的渐近置信序列?
- RQ3如何将渐近推断扩展至因果推断中的序列设置,特别是在非渐近界不存在的观测研究中?
- RQ4置信序列的宽度能否像经典CLT区间一样,根据估计方差自适应缩放,同时保持时间一致覆盖?
- RQ5是否存在收紧置信序列宽度的根本限制?所提出的构造在最小最大意义下是否最优?
主要发现
- 所提出的渐近置信序列在与经典中心极限定理相同的矩条件下具有普遍有效性,具备广泛适用性。
- 置信序列的宽度与高效影响函数估计标准差成比例,即使在非渐近设置下也实现方差自适应。
- 该方法在不依赖固定样本量的情况下实现时间一致覆盖,支持连续监测与自适应停止,且无第一类错误膨胀。
- 其核心组成部分——正态混合边界、近似误差率和标准误估计——由于强不变性原理的根本限制而不可改进。
- 该框架使得在序列监测下对观测研究和随机化研究中的平均处理效应进行因果推断成为可能,而这些场景下非渐近界无法获得。
- 该方法继承了Robbins正态混合边界的最优性以及样本均值被布朗运动几乎必然近似的性质,因此在序列设置下为最小最大最优。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。