[Paper Review] A Meta-Learning Method for Estimation of Causal Excursion Effects to Assess Time-Varying Moderation
This paper proposes a meta-learner framework using debiased/orthogonal machine learning to estimate causal excursion effects in micro-randomized trials (MRTs), enabling efficient and robust assessment of time-varying treatment moderation. By leveraging supervised learning for nuisance parameter estimation and ensuring double robustness, the method reduces bias and improves efficiency compared to traditional weighted least squares approaches.
Advances in wearable technologies and health interventions delivered by smartphones have greatly increased the accessibility of mobile health (mHealth) interventions. Micro-randomized trials (MRTs) are designed to assess the effectiveness of the mHealth intervention and introduce a novel class of causal estimands called "causal excursion effects." These estimands enable the evaluation of how intervention effects change over time and are influenced by individual characteristics or context. Existing methods for analyzing causal excursion effects assume known randomization probabilities, complete observations, and a linear nuisance function with prespecified features of the high dimensional observed history. However, in complex mobile systems, these assumptions often fall short: randomization probabilities can be uncertain, observations may be incomplete, and the granularity of mHealth data makes linear modeling difficult. To address this issue, we propose a flexible and doubly robust inferential procedure, called "DR-WCLS," for estimating causal excursion effects from a meta-learner perspective. We present the bidirectional asymptotic properties of the proposed estimators and compare them with existing methods both theoretically and through extensive simulations. The results show a consistent and more efficient estimate, even with missing observations or uncertain treatment randomization probabilities. Finally, the practical utility of the proposed methods is demonstrated by analyzing data from a multiinstitution cohort of first-year medical residents in the United States (NeCamp et al., 2020).
Motivation & Objective
- To address the challenge of model misspecification in causal excursion effect estimation from MRTs due to pre-specified, hand-crafted features of high-dimensional history.
- To improve estimation efficiency and robustness by replacing parametric working models with machine learning for nuisance parameters.
- To develop a doubly robust, asymptotically normal estimator for causal excursion effects using meta-learner principles and debiased estimation.
- To enable data-driven, flexible modeling of time-varying treatment moderation in mobile health interventions.
Proposed method
- The method employs a meta-learner framework that treats the estimation of nuisance parameters (e.g., outcome and treatment propensity models) as supervised learning problems.
- It uses Neyman-orthogonal estimating equations and cross-fitting to reduce regularization bias from machine learning algorithms.
- The core estimator is based on a weighted, centered least squares (WCLS) criterion, with nuisance components estimated via flexible ML models (e.g., random forests, BART, neural networks).
- The approach ensures double robustness: the final causal effect estimator is consistent if either the outcome regression or the propensity score model is correctly specified.
- The method applies debiased/orthogonal estimation to ensure asymptotic normality and valid inference under high-dimensional, time-varying covariates.
- The framework is extended to handle missing data, lagged effects, and binary outcomes through appropriate modeling adjustments.

Experimental results
Research questions
- RQ1Can machine learning be used effectively to estimate nuisance parameters in causal excursion effect estimation without introducing bias?
- RQ2How does the proposed meta-learner method improve estimation efficiency compared to traditional WCLS with pre-specified features?
- RQ3Does the method achieve double robustness in the presence of model misspecification for either the outcome or treatment model?
- RQ4What is the impact of using flexible ML models on the precision and coverage of confidence intervals for time-varying causal effects?
- RQ5How does the method perform in real-world mHealth data with complex, time-varying confounding?
Key findings
- The proposed doubly robust WCLS estimator achieves significant relative efficiency gains over standard WCLS, with narrower confidence intervals in both simulations and real data.
- The DR-WCLS method demonstrated improved coverage and reduced bias even when one of the nuisance models was misspecified, confirming double robustness.
- In the case study of first-year medical residents, mobile prompts had a positive causal effect on step count in early weeks, which diminished over time—suggesting habituation.
- The method successfully identified time-varying moderation effects, showing that intervention effectiveness depends on prior behavior, mood, and contextual factors.
- The simulation results confirmed asymptotic normality of the estimator and validated the theoretical properties under high-dimensional, time-varying covariates.
- The R-WCLS and DR-WCLS methods outperformed standard WCLS in terms of precision and power, especially in settings with complex, nonlinear relationships.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.