Skip to main content
QUICK REVIEW

[Paper Review] People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy

Shruthi Shekar, Pat Pataranutaporn|arXiv (Cornell University)|Aug 11, 2024
Artificial Intelligence in Healthcare and Education11 citations
TL;DR

The study shows non-experts can’t distinguish AI-generated medical replies from doctors, often trust AI as valid as or more trustworthy, especially when AI is labeled as high accuracy, raising safety concerns.

ABSTRACT

This paper presents a comprehensive analysis of how AI-generated medical responses are perceived and evaluated by non-experts. A total of 300 participants gave evaluations for medical responses that were either written by a medical doctor on an online healthcare platform, or generated by a large language model and labeled by physicians as having high or low accuracy. Results showed that participants could not effectively distinguish between AI-generated and Doctors' responses and demonstrated a preference for AI-generated responses, rating High Accuracy AI-generated responses as significantly more valid, trustworthy, and complete/satisfactory. Low Accuracy AI-generated responses on average performed very similar to Doctors' responses, if not more. Participants not only found these low-accuracy AI-generated responses to be valid, trustworthy, and complete/satisfactory but also indicated a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result of the response provided. This problematic reaction was comparable if not more to the reaction they displayed towards doctors' responses. This increased trust placed on inaccurate or inappropriate AI-generated medical advice can lead to misdiagnosis and harmful consequences for individuals seeking help. Further, participants were more trusting of High Accuracy AI-generated responses when told they were given by a doctor and experts rated AI-generated responses significantly higher when the source of the response was unknown. Both experts and non-experts exhibited bias, finding AI-generated responses to be more thorough and accurate than Doctors' responses but still valuing the involvement of a Doctor in the delivery of their medical advice. Ensuring AI systems are implemented with medical professionals should be the future of using AI for the delivery of medical advice.

Motivation & Objective

  • Assess whether laypeople can distinguish AI-generated medical responses from doctor-provided ones.
  • Evaluate perceptions of validity, trustworthiness, completeness, and user intentions toward AI vs. doctor responses.
  • Examine how knowledge of the response source influences perception and trust.
  • Explore whether expert evaluators exhibit biases when source is revealed or unknown.

Proposed method

  • Collect 150 AI-generated responses to HealthTap questions and have four physicians rate accuracy (Yes, Maybe, No) to classify High vs. Low Accuracy AI outputs.
  • Create a dataset with 30 Doctor responses, 30 High Accuracy AI responses, 30 Low Accuracy AI responses for 100 online participants.
  • Experiment 1: participants judge understanding and source, plus confidence in source; compare AI vs. Doctor.
  • Experiment 2: participants evaluate AI vs. Doctor without knowing exact source; measure validity, trust, completeness, and behavioral intentions.
  • Experiment 3: random label experiment to test bias by source descriptor (Doctor, AI, Doctor assisted by AI) on evaluations.
  • Additional physician evaluation under Blind/Non-Blind conditions to assess expert bias against AI when source is disclosed.
Figure 1: Visual summary of the dataset construction and pipeline of experiments discussed in this paper.
Figure 1: Visual summary of the dataset construction and pipeline of experiments discussed in this paper.

Experimental results

Research questions

  • RQ1Can participants distinguish AI-generated from Doctor-provided medical responses?
  • RQ2How do participants rate validity, trust, completeness, and behavioral intentions for AI vs. Doctor responses?
  • RQ3Does knowledge of the response source (Doctor vs. AI) influence perceptions and trust?
  • RQ4Do experts show biases when evaluating AI outputs with/without source disclosure?

Key findings

  • Participants could not reliably distinguish AI-generated from Doctor responses, with approximately 50% accuracy in source identification across types.
  • High Accuracy AI-generated responses were rated as significantly more valid, trustworthy, and complete/satisfactory than Doctor responses in Experiment 2.
  • Low Accuracy AI-generated responses performed similarly to Doctors’ responses across validity, trust, and completeness measures; in some cases they exceeded Doctor performance.
  • Participants generally trusted AI-generated responses when the source was unknown, but trust increased further when High Accuracy AI responses were labeled as from a Doctor.
  • Experts rated AI-generated responses higher overall when the source was unknown, but their ratings dropped significantly when informed the response came from AI.
  • The study highlights risks of misinformation: lay users may follow harmful AI advice if not supervised by a clinician.
Figure 2: Example Medical Questions by Category: Comparing Doctors’ and AI-Generated Responses
Figure 2: Example Medical Questions by Category: Comparing Doctors’ and AI-Generated Responses

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.