[Paper Review] People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy
The study shows non-experts can’t distinguish AI-generated medical replies from doctors, often trust AI as valid as or more trustworthy, especially when AI is labeled as high accuracy, raising safety concerns.
This paper presents a comprehensive analysis of how AI-generated medical responses are perceived and evaluated by non-experts. A total of 300 participants gave evaluations for medical responses that were either written by a medical doctor on an online healthcare platform, or generated by a large language model and labeled by physicians as having high or low accuracy. Results showed that participants could not effectively distinguish between AI-generated and Doctors' responses and demonstrated a preference for AI-generated responses, rating High Accuracy AI-generated responses as significantly more valid, trustworthy, and complete/satisfactory. Low Accuracy AI-generated responses on average performed very similar to Doctors' responses, if not more. Participants not only found these low-accuracy AI-generated responses to be valid, trustworthy, and complete/satisfactory but also indicated a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result of the response provided. This problematic reaction was comparable if not more to the reaction they displayed towards doctors' responses. This increased trust placed on inaccurate or inappropriate AI-generated medical advice can lead to misdiagnosis and harmful consequences for individuals seeking help. Further, participants were more trusting of High Accuracy AI-generated responses when told they were given by a doctor and experts rated AI-generated responses significantly higher when the source of the response was unknown. Both experts and non-experts exhibited bias, finding AI-generated responses to be more thorough and accurate than Doctors' responses but still valuing the involvement of a Doctor in the delivery of their medical advice. Ensuring AI systems are implemented with medical professionals should be the future of using AI for the delivery of medical advice.
Motivation & Objective
- Assess whether laypeople can distinguish AI-generated medical responses from doctor-provided ones.
- Evaluate perceptions of validity, trustworthiness, completeness, and user intentions toward AI vs. doctor responses.
- Examine how knowledge of the response source influences perception and trust.
- Explore whether expert evaluators exhibit biases when source is revealed or unknown.
Proposed method
- Collect 150 AI-generated responses to HealthTap questions and have four physicians rate accuracy (Yes, Maybe, No) to classify High vs. Low Accuracy AI outputs.
- Create a dataset with 30 Doctor responses, 30 High Accuracy AI responses, 30 Low Accuracy AI responses for 100 online participants.
- Experiment 1: participants judge understanding and source, plus confidence in source; compare AI vs. Doctor.
- Experiment 2: participants evaluate AI vs. Doctor without knowing exact source; measure validity, trust, completeness, and behavioral intentions.
- Experiment 3: random label experiment to test bias by source descriptor (Doctor, AI, Doctor assisted by AI) on evaluations.
- Additional physician evaluation under Blind/Non-Blind conditions to assess expert bias against AI when source is disclosed.

Experimental results
Research questions
- RQ1Can participants distinguish AI-generated from Doctor-provided medical responses?
- RQ2How do participants rate validity, trust, completeness, and behavioral intentions for AI vs. Doctor responses?
- RQ3Does knowledge of the response source (Doctor vs. AI) influence perceptions and trust?
- RQ4Do experts show biases when evaluating AI outputs with/without source disclosure?
Key findings
- Participants could not reliably distinguish AI-generated from Doctor responses, with approximately 50% accuracy in source identification across types.
- High Accuracy AI-generated responses were rated as significantly more valid, trustworthy, and complete/satisfactory than Doctor responses in Experiment 2.
- Low Accuracy AI-generated responses performed similarly to Doctors’ responses across validity, trust, and completeness measures; in some cases they exceeded Doctor performance.
- Participants generally trusted AI-generated responses when the source was unknown, but trust increased further when High Accuracy AI responses were labeled as from a Doctor.
- Experts rated AI-generated responses higher overall when the source was unknown, but their ratings dropped significantly when informed the response came from AI.
- The study highlights risks of misinformation: lay users may follow harmful AI advice if not supervised by a clinician.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.