[论文解读] Prediction under interventions: evaluation of counterfactual performance using longitudinal observational data
本文提出了一种新颖的方法,用于评估纵向观察数据中具有生存时间结局的干预模型的反事实预测性能。通过结合人工删失和逆概率加权,该方法生成模拟假设治疗策略的合成验证数据集,从而实现对校准度、区分度(c指数、AUCt)和Brier评分在干预下的严格评估——为现实世界场景(如器官分配)中的因果预测评估提供了稳健框架。
Predictions under interventions are estimates of what a person's risk of an outcome would be if they were to follow a particular treatment strategy, given their individual characteristics. Such predictions can give important input to medical decision making. However, evaluating predictive performance of interventional predictions is challenging. Standard ways of evaluating predictive performance do not apply when using observational data, because prediction under interventions involves obtaining predictions of the outcome under conditions that are different to those that are observed for a subset of individuals in the validation dataset. This work describes methods for evaluating counterfactual performance of predictions under interventions for time-to-event outcomes. This means we aim to assess how well predictions would match the validation data if all individuals had followed the treatment strategy under which predictions are made. We focus on counterfactual performance evaluation using longitudinal observational data, and under treatment strategies that involve sustaining a particular treatment regime over time. We introduce an estimation approach using artificial censoring and inverse probability weighting which involves creating a validation dataset that mimics the treatment strategy under which predictions are made. We extend measures of calibration, discrimination (c-index and cumulative/dynamic AUCt) and overall prediction error (Brier score) to allow assessment of counterfactual performance. The methods are evaluated using a simulation study, including scenarios in which the methods should detect poor performance. Applying our methods in the context of liver transplantation shows that our procedure allows quantification of the performance of predictions supporting crucial decisions on organ allocation.
研究动机与目标
- 为填补评估估计假设治疗干预下结果的模型预测性能方面的关键空白。
- 开发一种可推广的框架,用于评估具有时间至事件结局的纵向观察数据中的反事实性能。
- 克服现有验证方法(如子集方法)的局限性,后者因对观察到的治疗进行条件化而存在选择偏倚。
- 将标准性能度量(校准度、区分度(c指数、动态AUCt)和Brier评分)扩展至持续治疗策略下的反事实设置。
- 为评估现实临床场景(如肝移植)中的干预预测提供一种实用且经过验证的方法。
提出的方法
- 在验证数据集中对个体进行人工删失,以模拟假设治疗策略:对非干预者在时间零删失,对干预者在地标时间删失。
- 应用治疗和删失的逆概率加权(IPACW)以校正选择偏倚,对干预策略使用时不变权重,对非干预策略使用时变权重。
- 采用个体-地标方法创建独立的验证数据集 $V^1$ 和 $V^0$,分别代表在特定时间接受和未接受移植的个体。
- 在30天区间内采用分段方法估计时变IPACW权重,同时考虑移植状态、等待名单移除和行政删失。
- 引入来自原始验证数据Kaplan-Meier估计的额外行政删失权重 $G_c^{-1}(t)$,以校正随访结束时的偏倚。
- 在加权且人工删失的验证数据集上应用标准性能度量(c指数、AUCt、Brier评分、校准度),以评估反事实性能。

实验结果
研究问题
- RQ1当仅有观察数据可用时,如何对估计假设干预下结果的模型进行有意义的预测性能评估?
- RQ2在时间至事件设置中,用于验证干预预测的常用子集方法中,选择偏倚的影响是什么?
- RQ3人工删失与逆概率加权相结合,能否在持续治疗策略下产生有效的反事实性能估计?
- RQ4标准性能度量(校准度、区分度和总体误差)在纵向观察数据的反事实验证中如何适应?
- RQ5在已知场景下,该方法在检测模型性能不佳方面相较于现有方法的优越程度如何?
主要发现
- 所提出的方法通过人工删失和逆概率加权成功创建了有效的反事实验证数据集,实现了对假设干预下预测性能的无偏评估。
- 子集方法在实践中被广泛使用,但其易受选择偏倚影响,导致性能评估无效。
- 该方法能够可靠估计反事实条件下关键性能指标,包括c指数、动态AUCt和Brier评分。
- 时变IPACW权重对于非干预策略至关重要,因为在该策略下,因移植和等待名单移除导致的删失随时间而变化。
- 包含行政删失权重 $G_c^{-1}(t)$ 对于保持性能估计的有效性至关重要,特别是在随访时间固定的研究所中。
- 该方法成功应用于肝移植数据集,证明其在高风险临床决策场景(如器官分配)中的实用性。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。