[论文解读] Quantitative and Qualitative Evaluation of Explainable Deep Learning Methods for Ophthalmic Diagnosis
本研究基于临床医生反馈,评估了13种用于视网膜光学相干断层扫描(OCT)诊断深度学习模型的可解释AI方法。Deep Taylor显著性归因得分最高(中位数3.85/5),优于Guided Backpropagation和SHAP,表明模型可解释性对于眼科临床应用至关重要。
Background: The lack of explanations for the decisions made by algorithms such as deep learning has hampered their acceptance by the clinical community despite highly accurate results on multiple problems. Recently, attribution methods have emerged for explaining deep learning models, and they have been tested on medical imaging problems. The performance of attribution methods is compared on standard machine learning datasets and not on medical images. In this study, we perform a comparative analysis to determine the most suitable explainability method for retinal OCT diagnosis. Methods: A commonly used deep learning model known as Inception v3 was trained to diagnose 3 retinal diseases - choroidal neovascularization (CNV), diabetic macular edema (DME), and drusen. The explanations from 13 different attribution methods were rated by a panel of 14 clinicians for clinical significance. Feedback was obtained from the clinicians regarding the current and future scope of such methods. Results: An attribution method based on a Taylor series expansion, called Deep Taylor was rated the highest by clinicians with a median rating of 3.85/5. It was followed by two other attribution methods, Guided backpropagation and SHAP (SHapley Additive exPlanations). Conclusion: Explanations of deep learning models can make them more transparent for clinical diagnosis. This study compared different explanations methods in the context of retinal OCT diagnosis and found that the best performing method may not be the one considered best for other deep learning tasks. Overall, there was a high degree of acceptance from the clinicians surveyed in the study. Keywords: explainable AI, deep learning, machine learning, image processing, Optical coherence tomography, retina, Diabetic macular edema, Choroidal Neovascularization, Drusen
研究动机与目标
- 基于真实世界专家反馈,评估可解释AI方法在眼科诊断中的临床相关性。
- 识别在视网膜OCT影像深度学习模型中最具效果的归因方法。
- 评估可解释性技术在眼科医生中的实际效用和接受度。
- 在临床背景下,比较13种归因方法的定性与定量性能。
- 为面向医学影像应用的可解释AI系统未来发展提供指导。
提出的方法
- 对预训练的Inception v3模型进行微调,以分类三种视网膜疾病:脉络膜新生血管形成(CNV)、糖尿病性黄斑水肿(DME)和玻璃疣(drusen)的OCT扫描。
- 应用13种归因方法(包括Deep Taylor、Guided Backpropagation和SHAP)生成模型预测的显著性图。
- 由14名眼科医生对每种解释的临床重要性进行5分制评分,重点关注解剖学合理性与诊断相关性。
- 定量评估包括对临床医生评分的统计分析,以对归因方法进行排序。
- 收集临床医生的定性反馈,以评估可解释性工具在当前和未来临床工作流程中的效用。
- 本研究采用受控的、临床医生参与的评估框架,以确保临床实际相关性。
实验结果
研究问题
- RQ1哪种可解释AI方法能为视网膜OCT诊断中的深度学习模型生成最具临床意义的解释?
- RQ2临床医生如何评价不同归因方法的可解释性与解剖学合理性?
- RQ3临床医生在多大程度上信任并认为模型解释对诊断决策具有价值?
- RQ4在医学影像与标准基准上评估时,不同归因方法的性能是否存在差异?
- RQ5临床医生如何看待可解释AI在眼科中的实际与未来应用场景?
主要发现
- Deep Taylor归因获得临床医生评分最高的中位分(3.85/5),表明其具有强临床可解释性。
- Guided Backpropagation和SHAP分别位列第二和第三,评分虽强但略低于Deep Taylor。
- 临床医生普遍接受可解释性方法,强调其在提升临床信任度与采纳率方面的潜力。
- 研究发现,最佳表现方法因上下文而异,表明模型可解释性应在特定领域环境中进行评估。
- 临床医生反馈强调了解剖准确性与定位在建立诊断信心中的重要性。
- 临床医生明显偏好能突出与已知病理特征相对应的视网膜区域的解释方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。