Skip to main content
QUICK REVIEW

[论文解读] Deciphering Diagnoses: How Large Language Models Explanations Influence Clinical Decision Making

Dmitriy Umerenkov, Galina Zubkova|arXiv (Cornell University)|Oct 3, 2023
Machine Learning in HealthcareComputer Science被引用 3
一句话总结

本研究评估大型语言模型(LLMs)在将患者主诉与诊断关联时生成解释的方式,发现LLM生成的解释显著提高了医生对诊断的一致性认同度——然而在5%至30%的案例中存在错误,主要源于对症状-诊断关联的过度解读。研究揭示了LLMs在临床决策支持中的说服力,同时也强调了事实性错误带来的风险。

ABSTRACT

Clinical Decision Support Systems (CDSS) utilize evidence-based knowledge and patient data to offer real-time recommendations, with Large Language Models (LLMs) emerging as a promising tool to generate plain-text explanations for medical decisions. This study explores the effectiveness and reliability of LLMs in generating explanations for diagnoses based on patient complaints. Three experienced doctors evaluated LLM-generated explanations of the connection between patient complaints and doctor and model-assigned diagnoses across several stages. Experimental results demonstrated that LLM explanations significantly increased doctors' agreement rates with given diagnoses and highlighted potential errors in LLM outputs, ranging from 5% to 30%. The study underscores the potential and challenges of LLMs in healthcare and emphasizes the need for careful integration and evaluation to ensure patient safety and optimal clinical utility.

研究动机与目标

  • 评估LLM生成的解释对医生临床决策的影响。
  • 评估LLM解释在连接患者主诉与医学诊断方面的质量与事实准确性。
  • 识别LLM生成医学解释中的常见错误模式。
  • 研究解释对经验丰富的临床医生之间诊断一致率的影响。
  • 突出LLMs集成到临床决策支持系统(CDSS)中的风险与机遇。

提出的方法

  • 开展多阶段实验,三位经验丰富的医生评估LLM为患者主诉生成的解释及其对应诊断。
  • 使用LLM基于电子病历(EHR)风格的主诉生成连接症状与诊断的纯文本解释。
  • 在不同阶段收集并分析医生的评估:初始诊断、查看LLM解释后、以及对解释质量的后评估。
  • 对医生与诊断之间的一致性率进行定量分析,比较暴露于LLM解释前后的情况。
  • 对LLM错误进行定性分析,重点关注事实错误、过度解读及症状误归因。
  • 对比医生与LLM生成的解释,以评估其一致性与可靠性。
Figure 1: Inter-doctor agreement on what explanations are erroneous for GT predictions
Figure 1: Inter-doctor agreement on what explanations are erroneous for GT predictions

实验结果

研究问题

  • RQ1LLM生成的解释在多大程度上影响医生对给定诊断假设的一致性认同?
  • RQ2LLM在解释患者主诉与诊断关联时,常犯哪些类型的事实错误?
  • RQ3LLM解释的质量在不同症状-诊断组合之间如何变化?
  • RQ4与仅依赖临床判断相比,LLM解释在多大程度上影响了诊断决策?
  • RQ5医生如何看待LLM生成解释的清晰度与临床相关性?

主要发现

  • LLM生成的解释显著提高了医生对给定诊断的一致性认同率,显示出明显的说服力。
  • LLM的错误率在5%至30%之间,最常见的错误是将症状错误地归因于不正确的诊断。
  • 一种主要错误类型是模型倾向于‘虚构’不存在的症状-诊断关联,尤其在症状无法明确支持诊断时更为明显。
  • 医生常常忽略了LLM解释中的事实错误,而更关注症状是否与诊断相符。
  • 研究发现,同一组症状可能对应多个有效的诊断假设,而LLM的解释会影响医生接受哪一个诊断。
  • 尽管许多解释质量较高,但医生对解释质量的一致性评价较低,表明解释质量的评估具有主观性且依赖于具体情境。
Figure 2: Inter-doctor agreement on what explanations are erroneous for MODEL predictions
Figure 2: Inter-doctor agreement on what explanations are erroneous for MODEL predictions

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。