Skip to main content
QUICK REVIEW

[Paper Review] Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias

Yu He Ke, Rui Yang|arXiv (Cornell University)|Jan 26, 2024
Clinical Reasoning and Diagnostic Skills5 citations
TL;DR

The paper shows that a GPT-4 powered multi-agent framework can improve diagnostic accuracy in challenging cases by simulating clinical team dynamics and mitigating cognitive biases, raising top differential accuracy from 0% to 71.3% and final two differentials to 80%.

ABSTRACT

Background: Cognitive biases in clinical decision-making significantly contribute to errors in diagnosis and suboptimal patient outcomes. Addressing these biases presents a formidable challenge in the medical field. Objective: This study explores the role of large language models (LLMs) in mitigating these biases through the utilization of a multi-agent framework. We simulate the clinical decision-making processes through multi-agent conversation and evaluate its efficacy in improving diagnostic accuracy. Methods: A total of 16 published and unpublished case reports where cognitive biases have resulted in misdiagnoses were identified from the literature. In the multi-agent framework, we leveraged GPT-4 to facilitate interactions among four simulated agents to replicate clinical team dynamics. Each agent has a distinct role: 1) To make the final diagnosis after considering the discussions, 2) The devil's advocate and correct confirmation and anchoring bias, 3) The tutor and facilitator of the discussion to reduce premature closure bias, and 4) To record and summarize the findings. A total of 80 simulations were evaluated for the accuracy of initial diagnosis, top differential diagnosis and final two differential diagnoses. Results: In a total of 80 responses evaluating both initial and final diagnoses, the initial diagnosis had an accuracy of 0% (0/80), but following multi-agent discussions, the accuracy for the top differential diagnosis increased to 71.3% (57/80), and for the final two differential diagnoses, to 80.0% (64/80). Conclusions: The framework demonstrated an ability to re-evaluate and correct misconceptions, even in scenarios with misleading initial investigations. The LLM-driven multi-agent conversation framework shows promise in enhancing diagnostic accuracy in diagnostically challenging medical scenarios.

Motivation & Objective

  • Motivate the need to address cognitive biases in clinical diagnosis.
  • Investigate whether a multi-agent LLM framework can mitigate biases during diagnostic reasoning.
  • Evaluate improvements in diagnostic accuracy across initial, top differential, and final differential diagnoses.

Proposed method

  • Construct a four-agent framework with distinct roles using GPT-4: final-diagnosis agent, devil's advocate, tutor/facilitator, and recorder/summarizer.
  • Simulate clinical decision-making across 16 case reports with known cognitive biases.
  • Run 80 simulations to assess changes from initial diagnosis to top differential and final two differentials.
  • Evaluate accuracy metrics for initial, top differential, and final two differentials.

Experimental results

Research questions

  • RQ1Can a multi-agent LLM framework correct misleading initial investigations and cognitive biases in diagnostic reasoning?
  • RQ2What is the impact of agent roles (devil's advocate, tutor, recorder) on diagnostic accuracy?
  • RQ3How does diagnostic accuracy change from initial to top differential and to final two differentials under the multi-agent setup?
  • RQ4Are LLM-driven discussions capable of improving accuracy in diagnostically challenging scenarios?

Key findings

  • Initial diagnoses had 0% accuracy (0/80).
  • Top differential accuracy rose to 71.3% (57/80) after multi-agent discussions.
  • Final two differentials accuracy reached 80.0% (64/80).
  • The framework demonstrated re-evaluation and correction of misconceptions even with misleading initial data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.