[论文解读] User Trust on an Explainable AI-based Medical Diagnosis Support System
本研究评估了放射科医生在使用基于因果解释的可解释人工智能(XAI)系统进行胸部X光片诊断时的信任度,其中放射科医生可查看并修正模型生成的因果解释。尽管模型准确率达到74.1%,但仅有46.4%的放射科医生认同模型预测结果,自报信任度平均为3.2/5,表明即使具备可解释性,信任度仍有限,且存在对解释过度依赖的证据。
Recent research has supported that system explainability improves user trust and willingness to use medical AI for diagnostic support. In this paper, we use chest disease diagnosis based on X-Ray images as a case study to investigate user trust and reliance. Building off explainability, we propose a support system where users (radiologists) can view causal explanations for final decisions. After observing these causal explanations, users provided their opinions of the model predictions and could correct explanations if they did not agree. We measured user trust as the agreement between the model's and the radiologist's diagnosis as well as the radiologists' feedback on the model explanations. Additionally, they reported their trust in the system. We tested our model on the CXR-Eye dataset and it achieved an overall accuracy of 74.1%. However, the experts in our user study agreed with the model for only 46.4% of the cases, indicating the necessity of improving the trust. The self-reported trust score was 3.2 on a scale of 1.0 to 5.0, showing that the users tended to trust the model but the trust still needs to be enhanced.
研究动机与目标
- 调查XAI系统中因果解释对放射科医生在医学诊断中信任与依赖的影响。
- 在真实临床诊断环境中,同时测量观察到的信任(与模型预测及解释的一致性)和自报信任(用户评分)。
- 探究用户对模型解释进行修正的行为是否反映出过度依赖或自我依赖的倾向。
- 评估模型可解释性对不熟悉AI系统的放射科医生信任校准的影响。
- 识别阻碍AI在放射科临床整合的信任与可用性方面的差距。
提出的方法
- 基于CXR-Eye数据集开发用于胸部X光片诊断的可解释人工智能模型,生成最终预测的因果解释。
- 设计用户研究,放射科医生首先独立解读X光片,随后查看模型预测及其因果解释。
- 允许放射科医生在不认同解释时进行修改,将这些修改作为观察到的信任度指标。
- 通过放射科医生与模型预测的一致率(46.4%)及解释修改的Wilcoxon符号秩检验,测量观察到的信任度。
- 使用5分制李克特量表收集自报信任度,平均得分为3.2。
- 采用深度学习架构训练模型,并应用知识蒸馏原理以提升性能。
实验结果
研究问题
- RQ1提供因果解释在多大程度上影响放射科医生对AI诊断结果的观察到的信任?
- RQ2放射科医生在多大程度上认同模型的最终预测结果和解释?
- RQ3能够修正解释的行为是否揭示了过度依赖或自我依赖的行为?
- RQ4在基于XAI的诊断系统中,自报信任度与观察到的信任度相比如何?
- RQ5哪些因素限制了不熟悉AI系统的放射科医生的信任校准?
主要发现
- 该模型在CXR-Eye数据集上达到74.1%的测试准确率,表明其具备较强的性能表现。
- 放射科医生仅在46.4%的病例中认同模型的最终诊断,表明观察到的信任度较低。
- 用户与模型解释高度一致,经Wilcoxon符号秩检验确认,表明存在对解释的过度依赖。
- 自报信任度平均为3.2/5,表明信任度中等但不够坚定。
- 放射科医生普遍缺乏AI使用经验,可能阻碍其信任校准,即使提供了解释。
- 医学诊断中的不确定性以及AI系统中不明确的置信度信号,可能在即使具备因果解释的情况下,仍导致信任度低下。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。