[Paper Review] Nonparametric Heterogeneous Treatment Effect Estimation in Repeated Cross Sectional Designs
This paper proposes a nonparametric estimator for heterogeneous treatment effects in repeated cross-sectional difference-in-differences designs, leveraging orthogonality to achieve fast convergence rates for treatment effect estimation even when nuisance components (e.g., outcome and propensity models) are estimated at slower rates. The method enables flexible, semiparametric estimation under conditional parallel trends, with strong empirical performance across simulations and real-world applications.
Identifying heterogeneity in a population's response to a health or policy intervention is crucial for evaluating and informing policy decisions. We propose a novel heterogeneous treatment effect estimator in the difference-in-differences design with repeated cross sectional data, where we observe different samples of a population at two time periods separated by the onset of a policy intervention, as well as samples of a population that serves as the control. Our estimator has orthogonality properties that enable fast rates on learning the treatment effect while allowing slower rates for estimating nuisance components. Our proposal shows promising empirical performance across a variety of simulation setups.
Motivation & Objective
- To address the challenge of estimating heterogeneous treatment effects in repeated cross-sectional difference-in-differences (DiD) designs without assuming parametric functional forms.
- To relax the global parallel trends assumption by conditioning on observed covariates, enabling more credible causal inference when subgroups exhibit distinct pre-trends.
- To develop a method that achieves fast convergence rates for treatment effect estimation while allowing slower rates for nuisance components such as outcome and propensity models.
- To ensure robustness and efficiency through orthogonality and double-robust estimation principles, particularly in high-dimensional or complex confounding settings.
Proposed method
- The method models the outcome as $ Y_i = b(X_i) + S_i \cdot \xi(X_i) + T_i \cdot \rho(X_i) + T_i S_i \cdot \tau(X_i) + \varepsilon_i $, where $ \tau(X_i) $ is the heterogeneous treatment effect function.
- It employs cross-fitting and orthogonal estimating equations to ensure that estimation of $ \tau(x) $ is robust to slow rates in estimating nuisance functions like $ b(x), \xi(x), \rho(x) $.
- The estimator uses a doubly robust augmented inverse probability weighting (AIPW) framework, combining outcome regression and propensity score estimation for efficiency and robustness.
- It applies nonparametric regression techniques (e.g., kernel or machine learning-based methods) to estimate the nuisance components and treatment effect function in a high-dimensional covariate space.
- The method ensures orthogonality by constructing estimating equations that are locally insensitive to perturbations in nuisance estimators, enabling asymptotic normality under weak regularity conditions.
- A cross-fitting procedure is used to avoid overfitting in nuisance estimation, with data split into folds to produce independent estimates for each component.
Experimental results
Research questions
- RQ1Can we estimate heterogeneous treatment effects in repeated cross-sectional DiD designs without assuming linear or parametric functional forms for the treatment effect?
- RQ2How can we relax the global parallel trends assumption by conditioning on observed covariates to improve validity in the presence of subgroup heterogeneity?
- RQ3What estimation strategy ensures fast convergence for the treatment effect function even when nuisance components are estimated at slower rates?
- RQ4How can we achieve double robustness and orthogonality in high-dimensional, nonparametric settings for causal inference in DiD frameworks?
Key findings
- The proposed estimator achieves a fast convergence rate for the heterogeneous treatment effect function $ \tau(x) $, even when nuisance components are estimated at slower rates, due to its orthogonality structure.
- Empirical results on simulated data show that the method outperforms standard linear interaction models in capturing complex, nonlinear treatment effect heterogeneity.
- In an application to Angrist and Kugler (2008) data, the method reveals a significantly larger treatment effect on individuals in rural areas compared to urban areas, suggesting meaningful heterogeneity.
- A placebo test using 1998 and 2000 as pre- and post-treatment periods shows negligible estimated treatment effects across all methods, supporting the validity of the approach.
- The AIPW-based average treatment effect estimator is semiparametrically efficient and robust to misspecification of either the outcome regression or propensity score model.
- The method maintains stability even when propensity scores are small, thanks to the use of cross-fitting and orthogonal estimating equations that reduce sensitivity to estimation error.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.