[Paper Review] The Optimal Dynamic Treatment Rule SuperLearner: Considerations, Performance, and Application
This paper proposes the Optimal Dynamic Treatment Rule SuperLearner (ODTR SuperLearner), an ensemble machine learning method that combines flexible and parametric algorithms to estimate individualized treatment rules that maximize expected outcomes. It demonstrates through simulations and a real-world application to a criminal justice RCT that the approach outperforms traditional parametric models in uncovering treatment effect heterogeneity, especially when using a full library of algorithms, continuous metalearners, and outcome-based risk functions.
The optimal dynamic treatment rule (ODTR) framework offers an approach for understanding which kinds of patients respond best to specific treatments -- in other words, treatment effect heterogeneity. Recently, there has been a proliferation of methods for estimating the ODTR. One such method is an extension of the SuperLearner algorithm -- an ensemble method to optimally combine candidate algorithms extensively used in prediction problems -- to ODTRs. Following the "causal roadmap," we causally and statistically define the ODTR and provide an introduction to estimating it using the ODTR SuperLearner. Additionally, we highlight practical choices when implementing the algorithm, including choice of candidate algorithms, metalearners to combine the candidates, and risk functions to select the best combination of algorithms. Using simulations, we illustrate how estimating the ODTR using this SuperLearner approach can uncover treatment effect heterogeneity more effectively than traditional approaches based on fitting a parametric regression of the outcome on the treatment, covariates and treatment-covariate interactions. We investigate the implications of choices in implementing an ODTR SuperLearner at various sample sizes. Our results show the advantages of: (1) including a combination of both flexible machine learning algorithms and simple parametric estimators in the library of candidate algorithms; (2) using an ensemble metalearner to combine candidates rather than selecting only the best-performing candidate; (3) using the mean outcome under the rule as a risk function. Finally, we apply the ODTR SuperLearner to the "Interventions" study, an ongoing randomized controlled trial, to identify which justice-involved adults with mental illness benefit most from cognitive behavioral therapy (CBT) to reduce criminal re-offending.
Motivation & Objective
- To develop a robust, flexible method for estimating optimal dynamic treatment rules (ODTRs) that account for treatment effect heterogeneity across diverse patient subgroups.
- To address limitations of traditional subgroup analyses and parametric models in capturing complex interactions between covariates and treatment effects.
- To guide precision health by identifying which patients benefit most from specific interventions, such as CBT for reducing recidivism among justice-involved individuals with mental illness.
- To provide practical guidance on implementing ODTR SuperLearner, including library composition, metalearner choice, and risk function selection.
- To validate the method using simulations and real-world data from the 'Interventions' RCT, demonstrating improved performance in estimating individualized treatment effects.
Proposed method
- The ODTR SuperLearner extends the SuperLearner algorithm to causal inference by estimating the optimal dynamic treatment rule through an ensemble of candidate estimators.
- It uses a two-stage approach: first estimating the blip function (treatment effect conditional on covariates), then combining these estimates via a metalearner to form the final rule.
- The method employs a library of candidate algorithms, including both parametric models (e.g., GLMs) and flexible machine learning methods (e.g., random forests, lasso), to enhance model adaptability.
- The metalearner combines predictions from all candidate algorithms using either a discrete (selection-only) or continuous (weighted combination) approach, with the latter showing improved performance.
- The risk function for selecting the optimal combination is based on the mean outcome under the estimated rule ($ R_{E[Y_d]} $), which better reflects the true performance of the treatment rule than mean squared error.
- The approach follows the causal roadmap, ensuring identification of the ODTR under the assumption of no unmeasured confounding and consistent estimation of counterfactual outcomes.
Experimental results
Research questions
- RQ1How does the ODTR SuperLearner compare to traditional parametric models in estimating treatment effect heterogeneity across diverse covariate structures?
- RQ2What is the impact of including both flexible machine learning and parametric models in the candidate algorithm library on ODTR estimation performance?
- RQ3Does using a continuous metalearner that combines predictions from multiple algorithms outperform selecting a single best-performing algorithm?
- RQ4How does the choice of risk function—mean squared error versus mean outcome under the rule—affect the accuracy of the estimated ODTR?
- RQ5Can the ODTR SuperLearner effectively identify subgroups that benefit most from cognitive behavioral therapy in a real-world randomized trial of justice-involved individuals with mental illness?
Key findings
- The ODTR SuperLearner significantly outperforms traditional parametric models in estimating treatment effect heterogeneity, particularly in complex data-generating processes with non-linear interactions.
- Including both flexible machine learning algorithms and simple parametric estimators in the candidate library leads to more accurate and robust ODTR estimation compared to using only one type.
- Using a continuous metalearner that combines predictions from multiple algorithms yields better performance than selecting only the single best-performing candidate algorithm.
- The mean outcome under the rule ($ R_{E[Y_d]} $) as a risk function produces more accurate ODTR estimates than mean squared error, as it directly optimizes the expected outcome of the treatment rule.
- In the 'Interventions' RCT, the ODTR SuperLearner identified that CBT reduces re-arrest probability among individuals with low substance use, suggesting a differential treatment effect by this covariate.
- The method achieved a 78% match rate in assigning the true optimal treatment in simulation DGP 1 and 75% in DGP 2, demonstrating strong empirical performance across diverse scenarios.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.