[论文解读] Off-Policy Estimation of Long-Term Average Outcomes With Applications to Mobile Health
本文提出了一种基于不同行为策略收集的历史数据,对移动健康(mHealth)干预中长期平均结果进行离策略估计的方法。该方法引入了一种双重稳健估计器及其置信区间,用于评估预先指定的mHealth策略(如基于位置或始终治疗的策略)的性能,使用来自微随机化试验(MRT)的数据,表明基于位置的策略相比‘不作为’策略可使30分钟步数提高约22%。
Due to the recent advancements in wearables and sensing technology, health scientists are increasingly developing mobile health (mHealth) interventions. In mHealth interventions, mobile devices are used to deliver treatment to individuals as they go about their daily lives. These treatments are generally designed to impact a near time, proximal outcome such as stress or physical activity. The mHealth intervention policies, often called just-in-time adaptive interventions, are decision rules that map an individual’s current state (e.g., individual’s past behaviors as well as current observations of time, location, social activity, stress, and urges to smoke) to a particular treatment at each of many time points. The vast majority of current mHealth interventions deploy expert-derived policies. In this article, we provide an approach for conducting inference about the performance of one or more such policies using historical data collected under a possibly different policy. Our measure of performance is the average of proximal outcomes over a long time period should the particular mHealth policy be followed. We provide an estimator as well as confidence intervals. This work is motivated by HeartSteps, an mHealth physical activity intervention. Supplementary materials for this article are available online.
研究动机与目标
- 利用在不同行为策略下收集的数据,实现对mHealth策略长期平均近端结果的推断。
- 解决专家制定的mHealth干预策略缺乏数据驱动评估的问题。
- 在已知随机化概率的序列决策设置中,开发一种灵活且统计有效的策略评估方法。
- 通过支持对多个候选策略的性能比较,促进基于数据的即时自适应干预设计。
- 将微随机化试验(MRT)的效用从即时因果推断扩展到长期策略评估。
提出的方法
- 使用马尔可夫决策过程(MDP)框架对mHealth干预中的序列决策进行建模。
- 应用双重稳健估计器以估计长期平均奖励,结合结果回归与逆概率加权。
- 采用再生核希尔伯特空间(RKHS)与径向基函数核来建模相对值函数和Q函数。
- 推导估计器的渐近正态性,以构建策略性能的有效置信区间。
- 使用数据驱动的调参程序选择RKHS估计中的正则化参数。
- 在半参数框架中求解贝尔曼方程,以估计目标策略下的长期平均结果。
实验结果
研究问题
- RQ1我们能否利用在不同行为策略下收集的数据,准确估计目标mHealth策略的长期平均近端结果?
- RQ2当行为策略已知且为随机策略时,如何构建mHealth策略性能的有效置信区间?
- RQ3基于位置的治疗策略在长期体力活动结果方面是否优于‘不作为’或‘始终治疗’策略?
- RQ4在模型误设或小样本量条件下,该方法的表现如何?
- RQ5该方法能否扩展至二值近端结果或非平稳环境?
主要发现
- 基于位置的策略(π_location)下估计的平均30分钟步数为3.155,95%置信区间为[2.893, 3.417],表明相比‘不作为’策略具有统计显著的改善。
- ‘不作为’策略的估计平均奖励为2.962,差异(π_location - π_nothing)的95%置信区间为[-0.016, 0.402],表明步数可能提高约22%。
- ‘始终治疗’策略的估计平均奖励为3.127,差异(π_location - π_always)的95%置信区间为[-0.161, 0.217],表明两者无显著差异。
- 模拟结果显示该方法具有足够的覆盖概率,验证了所提置信区间的有效性。
- RKHS正则化参数的调参程序在有限样本中有效平衡了偏差与方差。
- 基于HeartSteps数据(n=37)的案例研究证明了该方法在具有长时程结果的真实mHealth数据中应用的可行性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。