[论文解读] Latent-state models for precision medicine
本文提出了一种部分可观察的马尔可夫决策过程(POMDP)模型,通过引入潜在健康状态,以估计在不规则、纵向观察数据中最优的动态治疗方案,尤其适用于双相情感障碍等慢性精神疾病。通过建模影响治疗反应的未观测患者状态,该方法能够在传统方法因缺乏时间对齐或马尔可夫假设而失效的情况下,实现个性化且具有临床意义的治疗策略估计。
Observational longitudinal studies are a common means to study treatment efficacy and safety in chronic mental illness. In many such studies, treatment changes may be initiated by either the patient or by their clinician and can thus vary widely across patients in their timing, number, and type. Indeed, in the observational longitudinal pathway of the STEP-BD study of bipolar depression, one of the motivations for this work, no two patients have the same treatment history even after coarsening clinic visits to a weekly time-scale. Estimation of an optimal treatment regime using such data is challenging as one cannot naively pool together patients with the same treatment history, as is required by methods based on inverse probability weighting, nor is it possible to apply backwards induction over the decision points, as is done in Q-learning and its variants. Thus, additional structure is needed to effectively pool information across patients and within a patient over time. Current scientific theory for many chronic mental illnesses maintains that a patient's disease status can be conceptualized as transitioning among a small number of discrete states. We use this theory to inform the construction of a partially observable Markov decision process model of patient health trajectories wherein observed health outcomes are dictated by a patient's latent health state. Using this model, we derive and evaluate estimators of an optimal treatment regime under two common paradigms for quantifying long-term patient health. The finite sample performance of the proposed estimator is demonstrated through a series of simulation experiments and application to the observational pathway of the STEP-BD study. We find that the proposed method provides high-quality estimates of an optimal treatment strategy in settings where existing approaches cannot be applied without ad hoc modifications.
研究动机与目标
- 解决在具有不规则、非马尔可夫性治疗史及频繁、可变治疗变化的观察性纵向研究中估计最优治疗方案的挑战。
- 将患者特异的潜在健康状态(如双相情感障碍中的未观测情绪状态)纳入治疗方案估计,以提高临床相关性与统计效率。
- 开发一种方法,使在无限时域设置下能够估计最优治疗方案,其中标准逆概率加权或Q-learning方法因缺乏时间对齐而无法适用。
- 提供一个理论基础坚实、临床可解释的框架,用于使用具有潜在状态动态的POMDP模型实现慢性精神疾病精准医疗。
- 通过模拟和对STEP-BD观察路径的应用,展示该方法的性能,证明其在复杂数据设置下优于现有方法。
提出的方法
- 使用部分可观察的马尔可夫决策过程(POMDP)对患者健康轨迹进行建模,其中观测结果依赖于未观测的潜在健康状态。
- 假设在给定当前潜在状态和观测协变量的条件下,患者健康的演化是马尔可夫性的,从而支持在不确定性下的顺序决策。
- 将信息状态定义为潜在状态与当前患者测量值的联合分布,该分布完全决定了最优治疗方案。
- 基于似然推断推导最优治疗方案的估计量,并在正则条件下建立其渐近性质。
- 将模型应用于估计在两种效用准则下的治疗方案:平均效用和折扣效用(γ = 0.95),使用来自STEP-BD研究的数据。
- 使用决策树可视化解释估计的最优治疗规则,分割基于抑郁、躁狂或其他情绪状态的估计概率。
实验结果
研究问题
- RQ1潜在状态POMDP模型是否能有效估计在具有不规则、非马尔可夫性治疗史的观察数据中最优治疗方案?
- RQ2将未观测的患者健康状态纳入模型,如何提升估计的动态治疗方案的质量与临床可解释性?
- RQ3在标准方法失效的情境下,所提估计量的有限样本性能与现有方法相比如何?
- RQ4估计的最优治疗方案在长期患者健康结局方面是否优于实际临床实践?
- RQ5不同的效用准则(平均效用 vs. 折扣效用)如何影响估计最优治疗方案的结构?
主要发现
- 所提出的基于POMDP的方法在STEP-BD观察路径中成功估计了最优治疗方案,而标准方法在无临时修改的情况下无法应用。
- 估计的最优治疗方案推荐仅使用情绪稳定剂或与抗抑郁药联合使用,避免单独使用抗抑郁药——这与临床指南一致,因单独使用抗抑郁药可能诱发躁狂。
- 对于平均效用,估计方案的值为1.81(95%置信区间:1.72–1.88),显著高于观察到的方案值1.66,且置信区间的下限高于观察值。
- 对于折扣效用(γ = 0.95),估计方案的值为33.76(95%置信区间:19.24–46.95),远高于观察值14.40。
- 该模型估计出具有临床意义的状态概率:例如,在‘抑郁’临床状态中,抑郁概率为85%;在‘躁狂’状态中,躁狂概率为55%。
- 在平均效用和折扣效用下,估计的最优治疗方案在定性上相似,抗抑郁药仅在抑郁可能性高或躁狂可能性低时被推荐。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。