[论文解读] Assessing the communication gap between AI models and healthcare professionals: explainability, utility and trust in AI-driven clinical decision-making
本研究通过癌症与新冠肺炎患者分诊的真实世界模拟(CORONET)评估了临床决策支持中的可解释人工智能,发现尽管解释能增强信任并减少自动化偏倚,但也可能引发确认偏倚和认知过载。研究提出了一套实用的机器学习可解释性评估框架,表明在模糊病例中,解释可改善决策,且对经验较少的临床医生具有支持作用。
This paper contributes with a pragmatic evaluation framework for explainable Machine Learning (ML) models for clinical decision support. The study revealed a more nuanced role for ML explanation models, when these are pragmatically embedded in the clinical context. Despite the general positive attitude of healthcare professionals (HCPs) towards explanations as a safety and trust mechanism, for a significant set of participants there were negative effects associated with confirmation bias, accentuating model over-reliance and increased effort to interact with the model. Also, contradicting one of its main intended functions, standard explanatory models showed limited ability to support a critical understanding of the limitations of the model. However, we found new significant positive effects which repositions the role of explanations within a clinical context: these include reduction of automation bias, addressing ambiguous clinical cases (cases where HCPs were not certain about their decision) and support of less experienced HCPs in the acquisition of new domain knowledge.
研究动机与目标
- 调查可解释机器学习模型在真实临床场景中对医疗专业人员决策影响的实际影响。
- 评估模型解释如何影响复杂、数据驱动决策环境中的信任度、可用性和临床表现。
- 识别解释在临床工作流程中产生的积极与消极影响,包括偏倚和认知负荷。
- 开发并验证一种评估临床环境中可解释人工智能实用性和可用性的框架,超越孤立的模型输出。
提出的方法
- 对23名医疗专业人员开展一项被试内用户研究,使用真实世界的临床决策支持工具(CORONET)进行癌症与新冠肺炎患者入院决策。
- 在10个模拟患者病例中,对比有解释(评分+颜色条)与无解释(评分+贡献特征)的模型表现。
- 采用一种实用的评估框架,将模型解释与临床实用性、心理模型及决策结果相联系。
- 收集关于感知信任度、可用性、认知努力和决策信心的定性与定量反馈。
- 使用领域专家设计的临床病例情景,确保临床一致性与真实性,尽管为人工构建。
- 分析解释对自动化偏倚、确认偏倚以及对模型局限性理解能力的影响。
实验结果
研究问题
- RQ1解释如何影响医疗专业人员在人工智能驱动的临床决策支持系统中的信任度与感知实用性?
- RQ2在临床决策中,解释的认知与行为效应是什么,包括确认偏倚与自动化偏倚?
- RQ3解释在多大程度上提升了临床医生检测模型错误或理解模型局限性能力?
- RQ4解释如何在临床判断模糊的病例中,或对经验较少的临床医生提供决策支持?
- RQ5当嵌入真实临床工作流程时,解释的实际效用如何,超越孤立的模型输出?
主要发现
- 解释减少了自动化偏倚,尤其当临床医生与模型推荐意见不一致时,促使他们进行更深入的反思。
- 对17%至35%的参与者而言,解释增加了认知努力,并导致了确认偏倚,可能强化了对模型的过度依赖。
- 57%的参与者即使在提供解释的情况下也未能检测到模型错误,表明其对模型局限性理解能力的提升有限。
- 解释显著改善了在临床判断模糊的病例中的决策质量,此时临床医生对自己的判断不确定。
- 经验较少的医疗专业人员报告称,解释有助于其获取新的临床领域知识。
- 尽管对解释作为信任机制持积极态度,但其实际效用取决于具体情境,并非始终有益。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。