[论文解读] Performance Deterioration of Deep Learning Models after Clinical Deployment: A Case Study with Auto-segmentation for Definitive Prostate Cancer Radiotherapy
本研究调查了在临床部署后,基于深度学习的前列腺癌放疗自动分割模型的性能退化问题。使用2006–2011年数据训练的UNet模型,在1,328名患者的2012–2022年数据上进行测试,结果显示自2015年后前列腺和直肠轮廓的DSC显著下降,其原因在于临床实践的演变,如水凝胶间隔物、层厚变化以及静脉造影剂的使用,凸显了持续监控和模型适应的必要性。
We evaluated the temporal performance of a deep learning (DL) based artificial intelligence (AI) model for auto segmentation in prostate radiotherapy, seeking to correlate its efficacy with changes in clinical landscapes. Our study involved 1328 prostate cancer patients who underwent definitive radiotherapy from January 2006 to August 2022 at the University of Texas Southwestern Medical Center. We trained a UNet based segmentation model on data from 2006 to 2011 and tested it on data from 2012 to 2022 to simulate real world clinical deployment. We measured the model performance using the Dice similarity coefficient (DSC), visualized the trends in contour quality using exponentially weighted moving average (EMA) curves. Additionally, we performed Wilcoxon Rank Sum Test to analyze the differences in DSC distributions across distinct periods, and multiple linear regression to investigate the impact of various clinical factors. The model exhibited peak performance in the initial phase (from 2012 to 2014) for segmenting the prostate, rectum, and bladder. However, we observed a notable decline in performance for the prostate and rectum after 2015, while bladder contour quality remained stable. Key factors that impacted the prostate contour quality included physician contouring styles, the use of various hydrogel spacer, CT scan slice thickness, MRI-guided contouring, and using intravenous (IV) contrast. Rectum contour quality was influenced by factors such as slice thickness, physician contouring styles, and the use of various hydrogel spacers. The bladder contour quality was primarily affected by using IV contrast. This study highlights the challenges in maintaining AI model performance consistency in a dynamic clinical setting. It underscores the need for continuous monitoring and updating of AI models to ensure their ongoing effectiveness and relevance in patient care.
研究动机与目标
- 评估深度学习自动分割模型在前列腺癌放疗临床部署后的长期性能稳定性。
- 识别导致人工智能轮廓勾画随时间性能下降的临床与影像因素。
- 评估16年期间前列腺、直肠和膀胱结构的Dice相似系数(DSC)的时间变化。
- 研究临床实践演变(如水凝胶间隔物、层厚和静脉造影剂使用)对模型性能的影响。
- 倡导建立持续的模型监控与再训练协议,以在动态临床环境中维持人工智能的有效性。
提出的方法
- 基于2006至2011年间1,328名前列腺癌患者扫描数据,训练了一个UNet深度学习模型,用于前列腺、直肠和膀胱的自动分割。
- 使用Dice相似系数(DSC)作为主要指标,在2012至2022年的测试数据上评估模型性能。
- 采用指数加权移动平均(EMA)曲线可视化不同时期轮廓质量的时间趋势。
- 应用Wilcoxon秩和检验检测不同时间区间内DSC分布的统计学显著差异。
- 使用多元线性回归量化临床因素(如层厚、水凝胶使用、静脉造影剂使用及医生轮廓绘制风格)对DSC值的影响。
- 通过在早期数据上进行训练并在逐步延后的时间数据上进行测试,模拟真实临床部署环境,反映临床实践的演变。
实验结果
研究问题
- RQ1在临床部署后,前列腺癌放疗的深度学习自动分割模型性能随时间如何变化?
- RQ2哪些临床因素与人工智能在前列腺、直肠和膀胱轮廓勾画中的性能下降关联最强?
- RQ3水凝胶间隔物的使用、CT层厚变化或静脉造影剂的使用是否显著影响模型随时间的准确性?
- RQ4医生轮廓绘制风格在不同时间段对模型DSC性能的影响程度如何?
- RQ5是否可以使用EMA曲线和统计检验可靠地追踪DSC的时间趋势,以检测模型性能退化?
主要发现
- 模型在2012至2014年表现出最佳性能,前列腺、直肠和膀胱结构的DSC值最高。
- 自2015年后,前列腺和直肠的DSC出现显著下降,而膀胱的DSC在时间上保持稳定。
- 前列腺轮廓质量最受医生轮廓绘制风格、水凝胶间隔物使用、CT层厚及静脉造影剂使用的影响。
- 直肠轮廓质量显著受层厚、医生轮廓绘制风格及水凝胶间隔物使用的影响。
- 膀胱轮廓质量主要受静脉(IV)造影剂使用的影响,其他方面无显著时间趋势。
- 本研究表明,临床实践的演变——而不仅仅是数据分布的变化——可导致已部署人工智能模型的性能退化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。