Skip to main content
QUICK REVIEW

[Paper Review] A comparative study of artificial intelligence and human doctors for the purpose of triage and diagnosis

Salman Razzaki, Adam Baker|arXiv (Cornell University)|Jun 27, 2018
Clinical Reasoning and Diagnostic Skills1 references44 citations
TL;DR

The study prospectively validates an AI triage and diagnostic system against human doctors using realistic vignettes, showing AI performance comparable to doctors and generally safer triage recommendations.

ABSTRACT

Online symptom checkers have significant potential to improve patient care, however their reliability and accuracy remain variable. We hypothesised that an artificial intelligence (AI) powered triage and diagnostic system would compare favourably with human doctors with respect to triage and diagnostic accuracy. We performed a prospective validation study of the accuracy and safety of an AI powered triage and diagnostic system. Identical cases were evaluated by both an AI system and human doctors. Differential diagnoses and triage outcomes were evaluated by an independent judge, who was blinded from knowing the source (AI system or human doctor) of the outcomes. Independently of these cases, vignettes from publicly available resources were also assessed to provide a benchmark to previous studies and the diagnostic component of the MRCGP exam. Overall we found that the Babylon AI powered Triage and Diagnostic System was able to identify the condition modelled by a clinical vignette with accuracy comparable to human doctors (in terms of precision and recall). In addition, we found that the triage advice recommended by the AI System was, on average, safer than that of human doctors, when compared to the ranges of acceptable triage provided by independent expert judges, with only a minimal reduction in appropriateness.

Motivation & Objective

  • Assess the diagnostic accuracy of an AI-powered triage and diagnostic system (Babylon) against human doctors.
  • Evaluate the safety and appropriateness of AI-driven triage recommendations.
  • Examine information-gathering and history-taking capabilities via a semi-naturalistic OSCE design.
  • Benchmark AI performance against publicly available case vignettes and established exam materials.

Proposed method

  • Use a semi-naturalistic role-play with mock consultations in an OSCE format.
  • Compare AI system outputs with independent blinded judges and multiple doctors.
  • Evaluate differential diagnoses and triage actions using recall, precision, and F1 metrics.
  • Incorporate expert qualitative ratings of differential quality and triage safety.
  • Test sensitivity of AI by varying internal thresholds to simulate doctor-type behaviors.

Experimental results

Research questions

  • RQ1Can an AI-powered triage and diagnostic system identify the condition modelled by a vignette with accuracy comparable to human doctors (precision and recall)?
  • RQ2Are AI-generated triage recommendations as safe as or safer than those provided by human doctors within independent judge thresholds?
  • RQ3How does AI performance fare against expert-rated differential quality and established exam benchmarks?
  • RQ4What is the impact of adjusting internal thresholds on the AI system’s recall vs. precision relative to doctors?
  • RQ5Do AI outputs generalize to publicly available vignette benchmarks (Semigran 2015, MRCGP AKT/CSA)?

Key findings

  • AI system achieves recall and precision comparable to doctors across vignettes (Babylon AI recall 80.0%, precision 44.4%, F1 57.1%).
  • Average doctor recall: 83.9%, precision: 43.6%, F1: 57.0% across seven doctors.
  • AI triage safety (97.0%) exceeds doctors (average 93.1%), with similar or slightly lower appropriateness (AI 90.0% vs doctors 90.5%).
  • Expert judge rated AI differential quality as comparable to doctors (83.0% and 83.0%–83.0%–? in different panels); GP-panel results varied, with AI sometimes rated lower depending on evaluator.
  • AI performance on Semigran 2015 vignettes: AI top-1 recall 70.0% and top-3 recall 96.7% versus doctors’ 75.3% and 90.3%.
  • AKT/CSA benchmarks showed AI top-3 inclusion of modelled disease in 86.7% (AKT) and 75.0% (CSA).

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.