[Paper Review] Double Machine Learning based Program Evaluation under Unconfoundedness
This paper proposes the normalised DR-learner (NDR-learner) to stabilize individualised treatment effect estimates in Double Machine Learning (DML)-based program evaluation under unconfoundedness. By applying individualised normalisation to inverse probability weights, the NDR-learner reduces extreme estimates from the standard DR-learner, improving robustness in finite samples while maintaining the flexibility and double-robustness of DML for estimating average, heterogeneous, and optimal treatment effects across Swiss Active Labour Market Policy programs.
This paper reviews, applies and extends recently proposed methods based on Double Machine Learning (DML) with a focus on program evaluation under unconfoundedness. DML based methods leverage flexible prediction models to adjust for confounding variables in the estimation of (i) standard average effects, (ii) different forms of heterogeneous effects, and (iii) optimal treatment assignment rules. An evaluation of multiple programs of the Swiss Active Labour Market Policy illustrates how DML based methods enable a comprehensive program evaluation. Motivated by extreme individualised treatment effect estimates of the DR-learner, we propose the normalised DR-learner (NDR-learner) to address this issue. The NDR-learner acknowledges that individualised effect estimates can be stabilised by an individualised normalisation of inverse probability weights.
Motivation & Objective
- To address finite-sample instability in individualised treatment effect estimates from the DR-learner in causal machine learning applications.
- To develop a method that stabilises inverse probability weights through individualised normalisation to improve estimation reliability.
- To demonstrate the practical utility of DML-based methods for comprehensive program evaluation in real-world policy settings.
- To provide a unified framework for estimating average effects, heterogeneous effects, and optimal treatment rules using flexible machine learning models.
- To evaluate the Swiss Active Labour Market Policy using DML methods and assess the role of caseworker characteristics in treatment effect variation.
Proposed method
- The NDR-learner applies individualised normalisation to inverse probability weights in the doubly robust score to stabilise individualised treatment effect estimates.
- It builds on the DR-learner framework, which uses machine learning to estimate potential outcomes and treatment probabilities for double-robust estimation.
- The method leverages the same doubly robust score for multiple estimands—average effects, heterogeneous effects, and optimal treatment rules—enabling computational and conceptual synergy.
- Estimation proceeds via cross-fitting to ensure valid inference, with t-tests and OLS used post-adjustment for causal parameters.
- The approach is applied to four Swiss ALMP programs using a standard dataset, with model selection and confounding adjustment handled via flexible machine learning.
- Optimal treatment assignment rules are estimated using a policy learning procedure based on the NDR-learner's output.
Experimental results
Research questions
- RQ1Can the DR-learner produce extreme individualised treatment effect estimates in finite samples, and if so, what causes this instability?
- RQ2Does individualised normalisation of inverse probability weights improve the stability and reliability of individualised treatment effect estimates?
- RQ3How do DML-based methods compare to traditional econometric approaches in estimating average, heterogeneous, and optimal treatment effects in program evaluation?
- RQ4What is the role of caseworker characteristics in explaining variation in individualised treatment effects in the Swiss ALMP context?
- RQ5To what extent do cross-validated policy rules for optimal treatment assignment agree with full-sample rules in this setting?
Key findings
- The DR-learner produced extreme individualised treatment effect estimates, with some values falling outside the theoretically possible range, indicating finite-sample instability.
- The NDR-learner successfully stabilised individualised treatment effect estimates by applying individualised normalisation to inverse probability weights, reducing extreme values.
- In the classification analysis, caseworker-related variables showed negligible influence on treatment effect variation, with all such coefficients in the lower half of the table.
- The optimal treatment assignment rules estimated via decision trees showed moderate overlap between cross-validated and full-sample policies, with agreement rates above 70% in both five- and 16-covariate models.
- The joint and marginal distributions of IATEs from the DR-learner and NDR-learner revealed that the latter produced more concentrated and plausible effect estimates.
- The study confirms that DML-based methods enable a comprehensive, computationally efficient evaluation of multiple programs under unconfoundedness, with strong statistical properties and reuse of the doubly robust score across estimands.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.