[Paper Review] Model-assisted analyses of cluster-randomized experiments
This paper proposes a design-based, model-assisted framework for analyzing cluster-randomized experiments, comparing regression estimators using individual-level, cluster-total, and cluster-average data. It demonstrates that covariate-adjusted regression on cluster totals yields the most efficient and consistent estimators, while robust standard errors provide conservative but reliable inference, even under model misspecification.
Cluster-randomized experiments are widely used due to their logistical convenience and policy relevance. To analyze them properly, we must address the fact that the treatment is assigned at the cluster level instead of the individual level. Standard analytic strategies are regressions based on individual data, cluster averages, and cluster totals, which differ when the cluster sizes vary. These methods are often motivated by models with strong and unverifiable assumptions, and the choice among them can be subjective. Without any outcome modeling assumption, we evaluate these regression estimators and the associated robust standard errors from a design-based perspective where only the treatment assignment itself is random and controlled by the experimenter. We demonstrate that regression based on cluster averages targets a weighted average treatment effect, regression based on individual data is suboptimal in terms of efficiency, and regression based on cluster totals is consistent and more efficient with a large number of clusters. We highlight the critical role of covariates in improving estimation efficiency, and illustrate the efficiency gain via both simulation studies and data analysis. Moreover, we show that the robust standard errors are convenient approximations to the true asymptotic standard errors under the design-based perspective. Our theory holds even when the outcome models are misspecified, so it is model-assisted rather than model-based. We also extend the theory to a wider class of weighted average treatment effects.
Motivation & Objective
- To address the inefficiency and model-dependency of standard regression methods in cluster-randomized experiments.
- To evaluate the design-based properties of regression estimators using individual data, cluster totals, and cluster averages.
- To quantify the efficiency-robustness trade-off in estimator selection under varying assumptions and covariate adjustments.
- To establish that robust standard errors are conservative approximations to true asymptotic standard errors in this context.
- To extend the framework to weighted average treatment effects and propose a new estimator based on weighted cluster averages.
Proposed method
- Adopt a design-based perspective where only treatment assignment is random, avoiding strong parametric assumptions on outcomes.
- Compare regression estimators on individual data, cluster totals, and cluster averages, with and without covariate adjustment.
- Use potential outcomes and design-based asymptotic theory to derive properties of estimators under randomization.
- Propose a weighted least squares estimator on cluster averages, with weights based on cluster size and covariates, to improve efficiency.
- Derive and evaluate robust standard errors (cluster-robust and heteroskedasticity-robust) as conservative approximations to true standard errors.
- Extend results to a broader class of weighted average treatment effects, including estimands that target cluster-level effects.

Experimental results
Research questions
- RQ1How do different regression estimators (on individual data, cluster totals, cluster averages) compare in terms of efficiency and consistency under design-based inference?
- RQ2What is the impact of covariate adjustment on estimator efficiency and robustness in cluster-randomized experiments?
- RQ3Are cluster-robust standard errors valid conservative approximations to true asymptotic standard errors in individual-level regressions?
- RQ4How does the choice of data level (individual vs. cluster) affect the efficiency-robustness trade-off in estimation?
- RQ5Can a new estimator based on weighted cluster averages outperform standard cluster-average regression in terms of efficiency and consistency?
Key findings
- Regression on cluster totals with covariate adjustment yields the most efficient and consistent estimator for the average treatment effect under large-sample design-based inference.
- Covariate-adjusted regression on individual data is suboptimal in efficiency but more robust to model misspecification and cluster size variability.
- Cluster-robust standard errors are conservative estimators of the true asymptotic standard errors in individual-level regressions, ensuring valid inference.
- Heteroskedasticity-robust standard errors are conservative in cluster-total regressions, extending Lin (2013) to cluster-randomized designs.
- The proposed weighted least squares estimator on cluster averages achieves higher efficiency than standard cluster-average regression by incorporating covariate information.
- Simulation and data analysis show that the covariate-adjusted cluster-total estimator reduces root mean squared error by up to 75% compared to unadjusted estimators.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.