[论文解读] Asymptotic Theory of Rerandomization in Treatment-Control Experiments
本文在不依赖高斯分布或可加性假设的前提下,为处理-对照实验中的重随机化发展了渐近理论,表明差异均值估计量具有非高斯的渐近分布——即标准正态变量与截断正态变量的线性组合。这使得能够构建精确的大样本置信区间,并表明与完全随机化相比,重随机化可降低抽样方差和四分位距。
Although complete randomization ensures covariate balance on average, the chance for observing significant differences between treatment and control covariate distributions increases with many covariates. Rerandomization discards randomizations that do not satisfy a predetermined covariate balance criterion, generally resulting in better covariate balance and more precise estimates of causal effects. Previous theory has derived finite sample theory for rerandomization under the assumptions of equal treatment group sizes, Gaussian covariate and outcome distributions, or additive causal effects, but not for the general sampling distribution of the difference-in-means estimator for the average causal effect. To supplement existing results, we develop asymptotic theory for rerandomization without these assumptions, which reveals a non-Gaussian asymptotic distribution for this estimator, specifically a linear combination of a Gaussian random variable and a truncated Gaussian random variable. This distribution follows because rerandomization affects only the projection of potential outcomes onto the covariate space but does not affect the corresponding orthogonal residuals. We also demonstrate that, compared to complete randomization, rerandomization reduces the asymptotic sampling variances and quantile ranges of the difference-in-means estimator. Moreover, our work allows the construction of accurate large-sample confidence intervals for the average causal effect, thereby revealing further advantages of rerandomization over complete randomization.
研究动机与目标
- 在不假设协变量、结果或可加性因果效应为高斯分布的前提下,发展重随机化的渐近理论。
- 在一般设定下,刻画重随机化下差异均值估计量的渐近分布。
- 表明与完全随机化相比,重随机化可降低估计量的渐近抽样方差和四分位距。
- 在无参数假设下,实现基于重随机化的平均因果效应的大样本置信区间的构建。
- 通过将潜在结果投影到协变量空间,阐明重随机化对潜在结果的几何影响。
提出的方法
- 使用几何框架将潜在结果分解为投影到协变量空间的分量与正交残差。
- 将差异均值估计量的渐近分布建模为标准正态变量与截断正态变量的线性组合。
- 应用马氏距离准则定义重随机化区域,限制协变量平衡较差的随机化方案。
- 采用条件渐近理论,即在协变量失衡在预设阈值范围内的条件下推导抽样分布。
- 通过协变量差异的标准化与正交分解,将重随机化的影响与残差变异分离。
- 依赖于重随机化仅影响结果在协变量空间上的投影,而不影响正交残差这一事实。
实验结果
研究问题
- RQ1在不作高斯分布或可加性假设的前提下,重随机化下差异均值估计量的渐近分布是什么?
- RQ2与完全随机化相比,重随机化如何影响差异均值估计量的抽样方差和四分位距?
- RQ3在无参数假设下,能否为重随机化下的平均因果效应构建准确的大样本置信区间?
- RQ4协变量空间在重随机化下的渐近分布形成中起什么作用?
- RQ5具体而言,重随机化的几何结构——特别是潜在结果的投影——如何影响推断?
主要发现
- 在无高斯或可加性假设下,重随机化下差异均值估计量的渐近分布是标准正态变量与截断正态变量的线性组合,而非高斯分布。
- 与完全随机化相比,重随机化降低了差异均值估计量的渐近抽样方差,从而提高了估计精度。
- 与完全随机化相比,重随机化下估计量的渐近四分位距更窄,表明抽样分布更集中。
- 通过将潜在结果分解为投影到协变量空间的分量与正交残差,推导出渐近分布,其中重随机化仅影响前者。
- 该理论支持在无高斯或可加性假设下,为重随机化下的平均因果效应构建有效的渐近置信区间。
- 核心洞见在于,重随机化仅改变结果在协变量空间上的投影,而正交残差保持不变,从而保持了渐近分布的结构。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。