Skip to main content
QUICK REVIEW

[论文解读] Model-assisted analyses of cluster-randomized experiments

Fangzhou Su, Peng Ding|arXiv (Cornell University)|Apr 9, 2021
Advanced Causal Inference Techniques参考文献 47被引用 4
一句话总结

本文提出了一种基于设计、模型辅助的框架,用于分析群组随机化实验,比较了使用个体层面数据、群组总和数据和群组平均数据的回归估计量。结果表明,对群组总和进行协变量调整的回归可得到最高效且一致的估计量,而稳健标准误即使在模型误设情况下也能提供保守但可靠的推断。

ABSTRACT

Cluster-randomized experiments are widely used due to their logistical convenience and policy relevance. To analyze them properly, we must address the fact that the treatment is assigned at the cluster level instead of the individual level. Standard analytic strategies are regressions based on individual data, cluster averages, and cluster totals, which differ when the cluster sizes vary. These methods are often motivated by models with strong and unverifiable assumptions, and the choice among them can be subjective. Without any outcome modeling assumption, we evaluate these regression estimators and the associated robust standard errors from a design-based perspective where only the treatment assignment itself is random and controlled by the experimenter. We demonstrate that regression based on cluster averages targets a weighted average treatment effect, regression based on individual data is suboptimal in terms of efficiency, and regression based on cluster totals is consistent and more efficient with a large number of clusters. We highlight the critical role of covariates in improving estimation efficiency, and illustrate the efficiency gain via both simulation studies and data analysis. Moreover, we show that the robust standard errors are convenient approximations to the true asymptotic standard errors under the design-based perspective. Our theory holds even when the outcome models are misspecified, so it is model-assisted rather than model-based. We also extend the theory to a wider class of weighted average treatment effects.

研究动机与目标

  • 解决标准回归方法在群组随机化实验中效率低下且依赖模型的问题。
  • 评估基于设计的个体数据、群组总和与群组平均数据的回归估计量的性质。
  • 量化在不同假设和协变量调整下估计量选择中的效率-稳健性权衡。
  • 证明稳健标准误在此情境下是对真实渐近标准误的保守近似。
  • 将框架扩展至加权平均处理效应,并提出一种基于加权群组平均的新估计量。

提出的方法

  • 采用基于设计的视角,仅假设处理分配是随机的,避免对结果变量施加强参数假设。
  • 比较在个体数据、群组总和与群组平均数据上使用与不使用协变量调整的回归估计量。
  • 利用潜在结果和基于设计的渐近理论,推导在随机化下估计量的性质。
  • 提出一种基于群组平均数据的加权最小二乘估计量,权重基于群组规模和协变量,以提高效率。
  • 推导并评估稳健标准误(群组稳健与异方差稳健)作为真实标准误的保守近似。
  • 将结果扩展至更广泛的加权平均处理效应类,包括针对群组层面效应的估计量。
(a) Simulation in Section 6.1 : comparing the estimators
(a) Simulation in Section 6.1 : comparing the estimators

实验结果

研究问题

  • RQ1在基于设计的推断下,不同回归估计量(基于个体数据、群组总和、群组平均)在效率和一致性方面如何比较?
  • RQ2协变量调整对群组随机化实验中估计量效率和稳健性有何影响?
  • RQ3群组稳健标准误是否是个体层面回归中真实渐近标准误的有效保守近似?
  • RQ4选择数据层级(个体 vs. 群组)如何影响估计中的效率-稳健性权衡?
  • RQ5基于加权群组平均的新估计量是否在效率和一致性方面优于标准群组平均回归?

主要发现

  • 在大样本基于设计的推断下,对群组总和进行协变量调整的回归可产生最高效且一致的平均处理效应估计量。
  • 对个体数据进行协变量调整的回归在效率上表现次优,但对模型误设和群组规模变异更具稳健性。
  • 在个体层面回归中,群组稳健标准误是对真实渐近标准误的保守估计,确保了有效的推断。
  • 在群组总和回归中,异方差稳健标准误也是保守的,扩展了Lin(2013)的结果至群组随机化设计。
  • 所提出的基于群组平均的加权最小二乘估计量通过整合协变量信息,其效率高于标准群组平均回归。
  • 模拟与数据分析表明,协变量调整的群组总和估计量相比未调整估计量,其均方根误差最高可降低75%。
(b) Simulation in Section 6.2 : noise covariate
(b) Simulation in Section 6.2 : noise covariate

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。