Skip to main content
QUICK REVIEW

[论文解读] Mind the Performance Gap: Examining Dataset Shift During Prospective Validation

Erkin Ötleş, Jeeheh Oh|arXiv (Cornell University)|Jul 23, 2021
Machine Learning in Healthcare参考文献 26被引用 9
一句话总结

本研究调查了用于预测医院相关感染的机器学习模型在回顾性验证与前瞻性验证之间性能差距的原因,发现基础设施变化——即数据提取和预处理流程的差异——是性能下降的主要驱动因素,而非患者人群或临床工作流程的时间性变化。模型的AUROC从回顾性验证的0.778下降至前瞻性验证的0.767,性能差距主要归因于数据管道不一致,而非临床变化。

ABSTRACT

Once integrated into clinical care, patient risk stratification models may perform worse compared to their retrospective performance. To date, it is widely accepted that performance will degrade over time due to changes in care processes and patient populations. However, the extent to which this occurs is poorly understood, in part because few researchers report prospective validation performance. In this study, we compare the 2020-2021 ('20-'21) prospective performance of a patient risk stratification model for predicting healthcare-associated infections to a 2019-2020 ('19-'20) retrospective validation of the same model. We define the difference in retrospective and prospective performance as the performance gap. We estimate how i) "temporal shift", i.e., changes in clinical workflows and patient populations, and ii) "infrastructure shift", i.e., changes in access, extraction and transformation of data, both contribute to the performance gap. Applied prospectively to 26,864 hospital encounters during a twelve-month period from July 2020 to June 2021, the model achieved an area under the receiver operating characteristic curve (AUROC) of 0.767 (95% confidence interval (CI): 0.737, 0.801) and a Brier score of 0.189 (95% CI: 0.186, 0.191). Prospective performance decreased slightly compared to '19-'20 retrospective performance, in which the model achieved an AUROC of 0.778 (95% CI: 0.744, 0.815) and a Brier score of 0.163 (95% CI: 0.161, 0.165). The resulting performance gap was primarily due to infrastructure shift and not temporal shift. So long as we continue to develop and validate models using data stored in large research data warehouses, we must consider differences in how and when data are accessed, measure how these differences may affect prospective performance, and work to mitigate those differences.

研究动机与目标

  • 调查临床环境中机器学习模型从回顾性验证转向前瞻性验证时性能退化的原因。
  • 区分时间性变化(患者人群和临床工作流程变化)与基础设施变化(数据访问、提取和转换流程差异)的影响。
  • 量化两种变化类型对真实世界医院相关感染患者风险分层模型性能差距的贡献。
  • 强调需要采用更贴近临床部署中实时数据访问的代表性回顾性数据管道。
  • 倡导将前瞻性验证作为临床部署前评估模型性能的关键步骤。

提出的方法

  • 作者比较了在单所学术医疗中心26,864例住院病例中,2019–2020年回顾性模型性能与2020–2021年前瞻性性能。
  • 将性能差距定义为回顾性与前瞻性验证之间AUROC和Brier评分的差异。
  • 将性能差距分解为两个组成部分:时间性变化(临床工作流程和患者人群变化)与基础设施变化(数据管道访问和预处理差异)。
  • 使用统计推断评估随时间的性能差异,包括逐月比较以及数据可用性和稳定性特征层面的分析。
  • 分析两个时期间特征分布(如住院位置、入院类型、药物医嘱)的变化,以评估基础设施和时间性变化的影响。
  • 使用AUROC和Brier评分评估模型校准与区分性能,并提供95%置信区间。

实验结果

研究问题

  • RQ1临床风险预测模型在回顾性与前瞻性验证之间的性能差距有多大?
  • RQ2性能差距在多大程度上由时间性变化(患者人群和临床工作流程变化)驱动,而非基础设施变化(数据访问和预处理差异)?
  • RQ3随时间推移,数据可用性和特征稳定性变化如何影响模型在真实世界部署中的性能?
  • RQ4由于疫情期间临床或操作变化,模型的区分性能是否随时间推移而恶化?
  • RQ5前瞻性验证能否帮助识别并量化回顾性评估未捕捉到的性能退化来源?

主要发现

  • 模型的AUROC从回顾性验证(2019–2020年)的0.778(95%置信区间:0.744–0.815)下降至前瞻性验证(2020–2021年)的0.767(95%置信区间:0.737–0.801),表明存在可测量的性能差距。
  • 性能差距主要由基础设施变化驱动——即数据提取和预处理流程的差异,而非患者人群或临床工作流程的时间性变化。
  • 时间性变化(包括因COVID-19疫情带来的变化,如新增患者科室、患者数量变化)并未恶化区分性能,甚至在某些月份(如2020年3月)略有改善。
  • 校准性能受到时间性变化的负面影响,表明尽管区分能力保持稳定,但模型在概率估计中的可靠性下降。
  • 药物给药数据随时间的稳定性高于生命体征数据,后者在回顾性数据中常被回溯记录,凸显了特定数据管道的不一致性。
  • 本研究强调,回顾性数据仓库可能无法准确反映实时数据的可及性,导致模型开发中性能估计过于乐观。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。