[Paper Review] Randomization-Based Causal Inference from Unbalanced 2^2 Split-Plot Designs
This paper develops a randomization-based causal inference method for unbalanced 2² split-plot designs, using a decomposition of potential outcomes to derive sampling variances as linear combinations of between- and within-block covariances. It demonstrates superior frequency coverage over model-based alternatives and establishes asymptotic connections to super-population inference.
Given two 2-level factors of interest, a 2^2 split-plot design} (a) takes each of the $2^2=4$ possible factorial combinations as a treatment, (b) identifies one factor as `whole-plot,' (c) divides the experimental units into blocks, and (d) assigns the treatments in such a way that all units within the same block receive the same level of the whole-plot factor. Assuming the potential outcomes framework, we propose in this paper a randomization-based estimation procedure for causal inference from 2^2 designs that are not necessarily balanced. Sampling variances of the point estimates are derived in closed form as linear combinations of the between- and within-block covariances of the potential outcomes. Results are compared to those under complete randomization as measures of design efficiency. Interval estimates are constructed based on conservative estimates of the sampling variances, and the frequency coverage properties evaluated via simulation. Asymptotic connections of the proposed approach to the model-based super-population inference are also established. Superiority over existing model-based alternatives is reported under a variety of settings for both binary and continuous outcomes.
Motivation & Objective
- To address the lack of randomization-based inference methods for unbalanced 2² split-plot designs in causal inference.
- To develop a variance estimation procedure that accounts for block structure and heterogeneity in experimental designs.
- To compare the efficiency of split-plot designs relative to complete randomization using empirical block-level heterogeneity.
- To reconcile finite-population randomization-based inference with model-based super-population inference via asymptotic arguments.
- To provide conservative interval estimates with validated coverage properties through simulation.
Proposed method
- Proposes a randomization-based estimation procedure under the potential outcomes framework for 2² split-plot designs.
- Derives sampling variances in closed form as linear combinations of between-block and within-block covariances of potential outcomes.
- Uses a decomposition of potential outcomes to link design efficiency to block-level heterogeneity.
- Applies projection matrix techniques to model the covariance structure under randomization.
- Constructs conservative confidence intervals using estimated sampling variances.
- Establishes asymptotic connections between finite-population randomization inference and linear mixed models via residual covariance limits.
Experimental results
Research questions
- RQ1How can randomization-based inference be extended to unbalanced 2² split-plot designs with block structure?
- RQ2What is the relative efficiency of split-plot designs compared to complete randomization, and how does it depend on block heterogeneity?
- RQ3How do the sampling variances of estimators in split-plot designs depend on within- and between-block covariances of potential outcomes?
- RQ4Can randomization-based inference achieve better frequency coverage than model-based methods under various outcome types (binary and continuous)?
- RQ5What is the asymptotic relationship between randomization-based inference and super-population model-based inference in this context?
Key findings
- The proposed randomization-based method achieves superior frequency coverage properties compared to model-based alternatives across diverse settings.
- Sampling variances are derived in closed form as linear combinations of between- and within-block covariances of potential outcomes.
- The method shows improved efficiency when block-level heterogeneity is high, linking design choice to empirical block structure.
- Conservative interval estimates based on the derived variances maintain good coverage in simulation studies.
- Asymptotic analysis supports the block-diagonal covariance structure assumed in linear mixed models, bridging finite-population and super-population inference.
- The approach outperforms existing model-based methods in terms of coverage accuracy for both binary and continuous outcomes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.