[Paper Review] Off-Policy Estimation of Long-Term Average Outcomes With Applications to Mobile Health
This paper proposes a method for off-policy estimation of long-term average outcomes in mobile health (mHealth) interventions using historical data from a different behavior policy. It introduces a doubly robust estimator with confidence intervals for evaluating the performance of pre-specified mHealth policies, such as those based on location or always treating, using data from a micro-randomized trial (MRT), demonstrating that location-based policies can increase 30-minute step counts by approximately 22% compared to 'do nothing' policies.
Due to the recent advancements in wearables and sensing technology, health scientists are increasingly developing mobile health (mHealth) interventions. In mHealth interventions, mobile devices are used to deliver treatment to individuals as they go about their daily lives. These treatments are generally designed to impact a near time, proximal outcome such as stress or physical activity. The mHealth intervention policies, often called just-in-time adaptive interventions, are decision rules that map an individual’s current state (e.g., individual’s past behaviors as well as current observations of time, location, social activity, stress, and urges to smoke) to a particular treatment at each of many time points. The vast majority of current mHealth interventions deploy expert-derived policies. In this article, we provide an approach for conducting inference about the performance of one or more such policies using historical data collected under a possibly different policy. Our measure of performance is the average of proximal outcomes over a long time period should the particular mHealth policy be followed. We provide an estimator as well as confidence intervals. This work is motivated by HeartSteps, an mHealth physical activity intervention. Supplementary materials for this article are available online.
Motivation & Objective
- To enable inference about the long-term average proximal outcomes of mHealth policies using data collected under a different behavior policy.
- To address the lack of data-driven evaluation for expert-derived mHealth intervention policies.
- To develop a flexible, statistically valid method for policy evaluation in sequential decision-making settings with known randomization probabilities.
- To support the design of data-based just-in-time adaptive interventions by enabling performance comparison of multiple candidate policies.
- To extend the utility of micro-randomized trials (MRTs) beyond immediate causal inference to long-term policy evaluation.
Proposed method
- Uses a Markov Decision Process (MDP) framework to model sequential decision-making in mHealth interventions.
- Applies a doubly robust estimator for long-term average rewards, combining outcome regression and inverse probability weighting.
- Employs a reproducing kernel Hilbert space (RKHS) with radial basis function kernel to model the relative value function and Q-functions.
- Derives asymptotic normality of the estimator to construct valid confidence intervals for policy performance.
- Uses a data-driven tuning procedure to select regularization parameters in the RKHS-based estimation.
- Solves the Bellman equation in a semi-parametric framework to estimate the long-term average outcome under a target policy.
Experimental results
Research questions
- RQ1Can we accurately estimate the long-term average proximal outcome of a target mHealth policy using data collected under a different behavior policy?
- RQ2How can we construct valid confidence intervals for the performance of mHealth policies when the behavior policy is known and stochastic?
- RQ3Does a location-based treatment policy outperform a 'do nothing' or 'always treat' policy in terms of long-term physical activity outcomes?
- RQ4How does the proposed method perform under model misspecification or small sample sizes?
- RQ5Can the method be extended to binary proximal outcomes or non-stationary environments?
Key findings
- The estimated average 30-minute step count under the location-based policy (π_location) was 3.155, with a 95% confidence interval of [2.893, 3.417], indicating a statistically significant improvement over the 'do nothing' policy.
- The 'do nothing' policy had an estimated average reward of 2.962, and the 95% confidence interval for the difference (π_location - π_nothing) was [-0.016, 0.402], suggesting a potential 22% increase in step count.
- The 'always treat' policy had an estimated average reward of 3.127, and the 95% confidence interval for the difference (π_location - π_always) was [-0.161, 0.217], indicating no significant difference between the two.
- The method achieved adequate coverage probability in simulations, validating the use of the proposed confidence intervals.
- The tuning procedure for RKHS regularization parameters was effective in balancing bias and variance in finite samples.
- The case study using HeartSteps data (n=37) demonstrated the feasibility of applying the method to real-world mHealth data with long-horizon outcomes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.