[论文解读] One Map Does Not Fit All: Evaluating Saliency Map Explanation on Multi-Modal Medical Images
本文提出 MSFI,一种新颖的评估指标,用于衡量多模态医学影像中显著图的性能,通过评估其突出对临床决策至关重要的模态特异性特征的能力。在 BraTS 和合成数据集上评估了 16 种显著图方法,研究发现大多数方法无法一致地定位模态特异性的重要特征,揭示了当前可解释人工智能在临床应用中的关键缺陷。
Being able to explain the prediction to clinical end-users is a necessity to leverage the power of AI models for clinical decision support. For medical images, saliency maps are the most common form of explanation. The maps highlight important features for AI model's prediction. Although many saliency map methods have been proposed, it is unknown how well they perform on explaining decisions on multi-modal medical images, where each modality/channel carries distinct clinical meanings of the same underlying biomedical phenomenon. Understanding such modality-dependent features is essential for clinical users' interpretation of AI decisions. To tackle this clinically important but technically ignored problem, we propose the MSFI (Modality-Specific Feature Importance) metric to examine whether saliency maps can highlight modality-specific important features. MSFI encodes the clinical requirements on modality prioritization and modality-specific feature localization. Our evaluations on 16 commonly used saliency map methods, including a clinician user study, show that although most saliency map methods captured modality importance information in general, most of them failed to highlight modality-specific important features consistently and precisely. The evaluation results guide the choices of saliency map methods and provide insights to propose new ones targeting clinical applications.
研究动机与目标
- 解决缺乏评估多模态医学影像中显著图是否反映临床有意义的模态特异性特征重要性的评估框架的问题。
- 开发一种与临床推理一致的指标,同时整合模态优先级和各模态下特征的精确定位。
- 使用基于临床的指标,在真实和合成的多模态 MRI 数据上评估 16 种现有显著图方法的性能。
- 通过识别最符合临床可解释性要求的方法,指导临床部署中显著图方法的选择。
- 为未来开发整合临床先验知识和模型表征保真度的显著图方法提供指导。
提出的方法
- 提出 MSFI(模态特异性特征重要性)作为复合指标,结合模态重要性(MI)和模态特异性特征定位。
- 使用从真实临床标注和模型预测中计算的 Shapley 值定义模态重要性,反映临床对模态的优先排序。
- 使用分割掩码计算模态特异性特征重要性,确保空间定位与临床相关性一致。
- 将 MSFI 分数归一化至 [0,1] 范围,以实现不同方法和数据集之间的比较,支持定量基准测试。
- 开展临床医生用户研究,将该指标与专家对 3D 显著图的评分进行对比,验证其与临床判断的相关性。
- 生成具有已知真实模态重要性的合成多模态 MRI 数据,以测试方法的鲁棒性和泛化能力。
实验结果
研究问题
- RQ1现有显著图方法是否能够准确突出多模态医学图像中的模态特异性特征,以满足临床推理需求?
- RQ2显著图方法在多大程度上与临床专家的模态优先级和特征定位偏好保持一致?
- RQ3所提出的 MSFI 指标在多大程度上与临床医生标注的显著图质量及临床知识相关?
- RQ4显著图方法在不同数据点上的表现是否一致,还是其临床可解释性存在高度变异性?
- RQ5MSFI 指标能否可靠地对显著图方法进行排序和选择,以用于临床决策支持系统中的部署?
主要发现
- 尽管捕捉到了一般模态重要性,大多数显著图方法仍无法一致且精确地突出模态特异性的重要特征。
- 在 BraTS 数据集上,所有方法的平均 MSFI 得分处于中等至偏低范围,且在单个测试案例间存在显著差异。
- GuidedBackProp 和 GuidedGradCAM 的平均 MSFI 得分最高,但仍表现出高方差,个别得分范围从 0 到 1。
- 临床医生用户研究显示,MSFI 得分与专家评分之间存在中等程度相关性,验证了该指标与临床判断的一致性。
- 在具有已知真实情况的合成数据集上(T1C 模态为主要模态),仅 GuidedBackProp 和 GuidedGradCAM 正确突出了 T1C 中的肿瘤,其余方法均失败。
- 本研究揭示,现有显著图方法可能因性能不稳定和模态特异性定位能力差,而无法可靠支持临床用户评估模型决策质量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。