[论文解读] Deep Learning for the Digital Pathologic Diagnosis of Cholangiocarcinoma and Hepatocellular Carcinoma: Evaluating the Impact of a Web-based Diagnostic Assistant
本研究评估了一款基于网络的深度学习诊断助手在全切片图像中区分肝细胞癌(HCC)与胆管细胞癌(CC)的性能。尽管在独立测试集上达到了84.2%的准确率,该助手并未提升整体病理医生的诊断准确率,且模型预测显著影响了病理医生的判断——在预测正确时提升表现,预测错误时则导致表现下降,凸显了在人工智能辅助病理学中锚定效应的风险。
While artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, the question of how best to incorporate these algorithms into clinical workflows remains relatively unexplored. We investigated how AI can affect pathologist performance on the task of differentiating between two subtypes of primary liver cancer, hepatocellular carcinoma (HCC) and cholangiocarcinoma (CC). We developed an AI diagnostic assistant using a deep learning model and evaluated its effect on the diagnostic performance of eleven pathologists with varying levels of expertise. Our deep learning model achieved an accuracy of 0.885 on an internal validation set of 26 slides and an accuracy of 0.842 on an independent test set of 80 slides. Despite having high accuracy on a hold out test set, the diagnostic assistant did not significantly improve performance across pathologists (p-value: 0.184, OR: 1.287 (95% CI 0.886, 1.871)). Model correctness was observed to significantly bias the pathologist decisions. When the model was correct, assistance significantly improved accuracy across all pathologist experience levels and for all case difficulty levels (p-value: < 0.001, OR: 4.289 (95% CI 2.360, 7.794)). When the model was incorrect, assistance significantly decreased accuracy across all 11 pathologists and for all case difficulty levels (p-value < 0.001, OR: 0.253 (95% CI 0.126, 0.507)). Our results highlight the challenges of translating AI models to the clinical setting, especially for difficult subspecialty tasks such as tumor classification. In particular, they suggest that incorrect model predictions could strongly bias an expert's diagnosis, an important factor to consider when designing medical AI-assistance systems.
研究动机与目标
- 评估基于网络的深度学习诊断助手是否能提升病理医生在区分HCC与CC方面的诊断准确率。
- 探究模型预测正确与否对病理医生决策的影响,特别是诊断偏倚的风险。
- 在真实临床工作流程模拟中,评估该助手在不同经验水平病理医生中的表现。
- 考察在亚专科病理学环境中部署AI决策支持工具的可行性和安全性。
提出的方法
- 使用来自癌症基因组图谱(The Cancer Genome Atlas)的70张H&E染色全切片图像(35例HCC,35例CC)对DenseNet-121卷积神经网络进行训练。
- 从肿瘤感兴趣区域提取图像块,用于模型的训练与验证,性能在内部验证集(26张切片)和独立外部测试集(80张切片)上进行评估。
- 通过云端部署的网络界面,允许病理医生上传选定的图像块以获取实时AI反馈,模拟临床第二意见工具的应用。
- 共11名病理医生(包括住院医师、非胃肠专科医生、胃肠专科医生及非胃肠(NOC)病理医生)参与实验,采用交叉设计对80张全切片图像进行判读,其中一半病例有辅助,另一半无辅助。
- 采用混合效应逻辑回归模型评估辅助作用及模型正确性对诊断准确率的影响,显著性通过Wald卡方检验进行评估。
- 在切片层面使用0.5的概率阈值评估模型性能,以分类HCC或CC。
实验结果
研究问题
- RQ1基于深度学习的诊断助手是否能提升病理医生在区分HCC与CC方面的诊断准确率?
- RQ2AI模型预测的正确性如何影响病理医生的诊断表现?
- RQ3该助手的影响是否因病理医生经验水平不同而有所差异?
- RQ4当AI助手预测错误时,其在多大程度上引入了诊断偏倚,特别是锚定效应?
主要发现
- 深度学习模型在内部验证集上达到88.5%的准确率,在包含80张全切片图像的独立外部测试集上达到84.2%的准确率。
- 辅助使用并未带来整体病理医生诊断准确率的统计学显著提升(p值:0.184,优势比:1.287,95%置信区间:0.886–1.871)。
- 当AI模型预测正确时,病理医生的准确率显著提高(p < 0.001,优势比:4.289,95%置信区间:2.360–7.794)。
- 当AI模型预测错误时,病理医生的准确率显著下降(p < 0.001,优势比:0.253,95%置信区间:0.126–0.507)。
- 模型输出带来的偏倚效应在所有病理医生经验水平和病例难度水平中均保持一致。
- 本研究揭示了锚定偏倚的严重风险:即使对经验丰富的病理医生,错误的AI预测也会产生强烈误导,从而危及人工智能辅助诊断的安全性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。