Skip to main content
QUICK REVIEW

[论文解读] Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias

Yu He Ke, Rui Yang|arXiv (Cornell University)|Jan 26, 2024
Clinical Reasoning and Diagnostic Skills被引用 5
一句话总结

本文表明一个由 GPT-4 驱动的多代理框架能够通过模拟临床团队动态和减轻认知偏差,在具有挑战性的病例中提高诊断准确性,将顶级鉴别诊断的准确率从 0% 提升到 71.3%,最终两项鉴别诊断达到 80%。

ABSTRACT

Background: Cognitive biases in clinical decision-making significantly contribute to errors in diagnosis and suboptimal patient outcomes. Addressing these biases presents a formidable challenge in the medical field. Objective: This study explores the role of large language models (LLMs) in mitigating these biases through the utilization of a multi-agent framework. We simulate the clinical decision-making processes through multi-agent conversation and evaluate its efficacy in improving diagnostic accuracy. Methods: A total of 16 published and unpublished case reports where cognitive biases have resulted in misdiagnoses were identified from the literature. In the multi-agent framework, we leveraged GPT-4 to facilitate interactions among four simulated agents to replicate clinical team dynamics. Each agent has a distinct role: 1) To make the final diagnosis after considering the discussions, 2) The devil's advocate and correct confirmation and anchoring bias, 3) The tutor and facilitator of the discussion to reduce premature closure bias, and 4) To record and summarize the findings. A total of 80 simulations were evaluated for the accuracy of initial diagnosis, top differential diagnosis and final two differential diagnoses. Results: In a total of 80 responses evaluating both initial and final diagnoses, the initial diagnosis had an accuracy of 0% (0/80), but following multi-agent discussions, the accuracy for the top differential diagnosis increased to 71.3% (57/80), and for the final two differential diagnoses, to 80.0% (64/80). Conclusions: The framework demonstrated an ability to re-evaluate and correct misconceptions, even in scenarios with misleading initial investigations. The LLM-driven multi-agent conversation framework shows promise in enhancing diagnostic accuracy in diagnostically challenging medical scenarios.

研究动机与目标

  • 动机解决临床诊断中的认知偏差的必要性。
  • 研究多代理 LLM 框架在诊断推理过程中是否能够减轻偏差。
  • 评估在初始诊断、首要鉴别诊断和最终鉴别诊断中的诊断准确性改进。

提出的方法

  • 使用 GPT-4 构建一个具有不同角色的四代理框架:最终诊断代理、反对者、导师/促进者,以及记录员/摘要员。
  • 对16份具有已知认知偏差的病例报告进行临床决策模拟。
  • 运行 80 次仿真以评估从初始诊断到首要鉴别诊断以及最终两项鉴别诊断的变化。
  • 评估初始、首要鉴别诊断以及最终两项鉴别诊断的准确性指标。

实验结果

研究问题

  • RQ1多代理 LLM 框架是否能够纠正诊断推理中误导性的初始调查和认知偏差?
  • RQ2代理角色(反对者、导师、记录员)对诊断准确性的影响是什么?
  • RQ3在多代理设置下,诊断准确性从初始到首要鉴别诊断再到最终两项鉴别诊断的变化如何?
  • RQ4基于LLM的讨论在诊断挑战性情景中是否能够提高准确性?

主要发现

  • 初始诊断的准确率为 0%(0/80)。
  • 经多代理讨论后,首要鉴别诊断的准确率上升至 71.3%(57/80)。
  • 最终两项鉴别诊断的准确率达到 80.0%(64/80)。
  • 该框架证明了即使初始数据具有误导性,也能重新评估并纠正误解。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。