Skip to main content
QUICK REVIEW

[论文解读] Three Applications of Conformal Prediction for Rating Breast Density in Mammography

Charles Lu, Ken Chang|arXiv (Cornell University)|Jun 23, 2022
AI in cancer detection被引用 5
一句话总结

本文将合规预测应用于乳腺X线摄影中的乳腺密度自动评估,展示了其在量化分布偏移检测的不确定性、通过选择性分类提升预测质量以及评估不同亚组公平性方面的实用性。结果表明,合规预测在提供理论保证的同时,显著提升了深度学习模型的可信度和临床可用性,性能与公平性均实现可测量的提升。

ABSTRACT

Breast cancer is the most common cancers and early detection from mammography screening is crucial in improving patient outcomes. Assessing mammographic breast density is clinically important as the denser breasts have higher risk and are more likely to occlude tumors. Manual assessment by experts is both time-consuming and subject to inter-rater variability. As such, there has been increased interest in the development of deep learning methods for mammographic breast density assessment. Despite deep learning having demonstrated impressive performance in several prediction tasks for applications in mammography, clinical deployment of deep learning systems in still relatively rare; historically, mammography Computer-Aided Diagnoses (CAD) have over-promised and failed to deliver. This is in part due to the inability to intuitively quantify uncertainty of the algorithm for the clinician, which would greatly enhance usability. Conformal prediction is well suited to increase reliably and trust in deep learning tools but they lack realistic evaluations on medical datasets. In this paper, we present a detailed analysis of three possible applications of conformal prediction applied to medical imaging tasks: distribution shift characterization, prediction quality improvement, and subgroup fairness analysis. Our results show the potential of distribution-free uncertainty quantification techniques to enhance trust on AI algorithms and expedite their translation to usage.

研究动机与目标

  • 为解决临床中因缺乏不确定性量化而导致对深度学习模型在乳腺X线摄影中乳腺密度评估可信度低的问题。
  • 评估合规预测作为无分布假设方法在真实医学影像数据集中不确定性量化的有效性。
  • 展示三个实际应用场景:检测分布偏移、通过选择性分类提升预测质量,以及分析亚组公平性。
  • 为临床医生提供可解释、理论基础坚实的不确定性估计,以支持临床决策。

提出的方法

  • 采用基于深度学习模型(EfficientNet-B3)的Softmax输出计算非一致性得分,应用合规预测进行乳腺密度分类。
  • 通过基于非一致性得分对测试样本进行排序,并在保留的验证集上进行校准,构建预测集。
  • 在分布偏移分析中,比较内部(DMIST)与外部(MGH)测试集的预测集大小,以检测性能下降。
  • 通过选择性分类策略,对高不确定性样本(集合大小=4)进行预测延迟,从而提升Cohen’s kappa等整体性能指标。
  • 通过在各亚组(年龄、种族、乳房大小、扫描仪类型)分别校准合规预测,开展公平性分析,以检测不确定性差异。
  • 使用p值(p < 0.001)评估不同亚组间预测集大小差异的统计显著性。

实验结果

研究问题

  • RQ1合规预测能否有效检测多中心乳腺X线摄影数据中的分布偏移?
  • RQ2通过合规预测实现的选择性分类在多大程度上可提升自动化乳腺密度预测的质量?
  • RQ3在年龄、种族、乳房大小和扫描仪类型等临床相关亚组中,是否存在可测量的不确定性差异?
  • RQ4在真实临床环境中,合规预测的不确定性量化与标准深度学习模型相比表现如何?

主要发现

  • 在α = 0.05时,合规预测将覆盖违反率降低至0.07,同时通过选择性分类将Cohen’s kappa从0.61提升至0.72,代价是延迟了56%的预测。
  • 外部测试集(MGH)的预测集大小显著高于内部测试集(DMIST),且全集预测(大小为4)更频繁,表明存在分布偏移。
  • 亚组分析显示,乳房较大的患者、拉美裔/西班牙裔患者、50岁以上患者以及使用B型扫描仪的患者,其平均预测集大小存在统计显著差异(p < 0.001)。
  • 在外部测试集中,致密型乳腺的平均集合大小最高,表明该类别在泛化时不确定性更高。
  • 合规预测框架成功识别出与已知混杂因素(如年龄和乳房体积)相关的临床相关不确定性模式。
  • 本研究证明,合规预测可提供可解释、理论基础坚实的不确定性估计,从而增强临床人工智能系统中的可信度与公平性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。