[Paper Review] Synthetic Interventions
This paper introduces Synthetic Interventions (SI), a causal inference framework that estimates $N \times D$ individual treatment effects using only two interventions per unit, leveraging a low-rank tensor factor model across units, interventions, and time. The method ensures finite-sample consistency and asymptotic normality even under unobserved confounders, validated via simulations and a large-scale e-commerce A/B test.
The synthetic controls (SC) methodology is a prominent tool for policy evaluation in panel data applications. Researchers commonly justify the SC framework with a low-rank matrix factor model that assumes the potential outcomes are described by low-dimensional unit and time specific latent factors. In the recent work of [Abadie '20], one of the pioneering authors of the SC method posed the question of how the SC framework can be extended to multiple treatments. This article offers one resolution to this open question that we call synthetic interventions (SI). Fundamental to the SI framework is a low-rank tensor factor model, which extends the matrix factor model by including a latent factorization over treatments. Under this model, we propose a generalization of the standard SC-based estimators. We prove the consistency for one instantiation of our approach and provide conditions under which it is asymptotically normal. Moreover, we conduct a representative simulation to study its prediction performance and revisit the canonical SC case study of [Abadie-Diamond-Hainmueller '10] on the impact of anti-tobacco legislations by exploring related questions not previously investigated.
Motivation & Objective
- To estimate $N \times D$ causal parameters—individual treatment effects—across heterogeneous units and multiple interventions.
- To reduce experimental cost by requiring only two interventions per unit, avoiding the need for $N \times D$ full randomized trials.
- To handle unobserved confounders by modeling them through low-dimensional latent factors in a tensor factor model.
- To provide finite-sample consistency and asymptotic normality for the estimator under mild regularity conditions.
- To enable data-efficient experimental design in settings like personalized policy evaluation or e-commerce optimization.
Proposed method
- Proposes a tensor factor model where potential outcomes are decomposed into low-rank components across units, interventions, and time points.
- Uses a modified principal component regression (PCR) estimator to predict counterfactual outcomes by learning weights from donor units under control.
- Employs a two-stage estimation: first, estimate donor weights via regression on pre-intervention data; second, predict post-intervention outcomes using these weights.
- Incorporates a confidence interval construction based on estimated error variance and weight norm, enabling inference under asymptotic normality.
- Handles unobserved confounding by assuming confounding effects are captured in the low-rank latent structure.
- Validates the method using synthetic data with known ground truth and a real-world A/B test on an e-commerce platform.
Experimental results
Research questions
- RQ1Can we estimate all $N \times D$ individual treatment effects with only two interventions per unit, without requiring $N \times D$ experiments?
- RQ2How can we ensure estimation consistency and valid inference when unobserved confounders affect intervention assignment?
- RQ3To what extent does the low-rank tensor factor model across units, interventions, and time enable information sharing and accurate prediction?
- RQ4What conditions ensure asymptotic normality of the estimator under the proposed model?
- RQ5How well does the method perform in practice compared to standard approaches, especially under limited data?
Key findings
- The SI estimator achieves finite-sample consistency under the proposed tensor factor model, even when unobserved confounders are present.
- Empirical coverage of 90% and 95% confidence intervals is close to nominal levels (88–90% and 94–95%) across all tested sample sizes $T_0 \in \{200, 400, 600, 800, 1000\}$, confirming valid inference.
- Interval lengths decrease with increasing $T_0$, but remain relatively large due to poorer prediction quality compared to standard PCR.
- The method successfully recovers causal parameters in a large-scale e-commerce A/B test, demonstrating practical utility in real-world settings.
- The estimator remains valid under unobserved confounding as long as the confounding structure is captured by the low-dimensional latent factors.
- The framework enables scalable causal inference in settings with high-dimensional interventions and heterogeneous units, reducing experimental burden significantly.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.