Skip to main content
QUICK REVIEW

[Paper Review] Putting ChatGPT's Medical Advice to the (Turing) Test

Oded Nov, Nina Singh|arXiv (Cornell University)|Jan 24, 2023
Digital Mental Health Interventions4 citations
TL;DR

This study evaluates whether ChatGPT can convincingly emulate human medical providers in patient communication by testing if laypeople can distinguish between AI-generated and real physician responses. Participants correctly identified response sources at 65.5% accuracy on average, indicating ChatGPT’s medical advice is weakly distinguishable from human responses, with moderate trust in chatbots for low-complexity health queries.

ABSTRACT

Objective: Assess the feasibility of using ChatGPT or a similar AI-based chatbot for patient-provider communication. Participants: A US representative sample of 430 study participants aged 18 and above. 53.2% of respondents analyzed were women; their average age was 47.1. Exposure: Ten representative non-administrative patient-provider interactions were extracted from the EHR. Patients' questions were placed in ChatGPT with a request for the chatbot to respond using approximately the same word count as the human provider's response. In the survey, each patient's question was followed by a provider- or ChatGPT-generated response. Participants were informed that five responses were provider-generated and five were chatbot-generated. Participants were asked, and incentivized financially, to correctly identify the response source. Participants were also asked about their trust in chatbots' functions in patient-provider communication, using a Likert scale of 1-5. Results: The correct classification of responses ranged between 49.0% to 85.7% for different questions. On average, chatbot responses were correctly identified 65.5% of the time, and provider responses were correctly distinguished 65.1% of the time. On average, responses toward patients' trust in chatbots' functions were weakly positive (mean Likert score: 3.4), with lower trust as the health-related complexity of the task in questions increased. Conclusions: ChatGPT responses to patient questions were weakly distinguishable from provider responses. Laypeople appear to trust the use of chatbots to answer lower risk health questions.

Motivation & Objective

  • To evaluate the feasibility of using AI chatbots like ChatGPT for patient-provider communication.
  • To assess whether laypeople can distinguish between ChatGPT-generated and human-generated medical responses.
  • To examine user trust in chatbots for handling health-related inquiries.
  • To explore how health-related complexity affects perceived trustworthiness of AI medical advice.

Proposed method

  • Ten real patient questions were extracted from electronic health records (EHRs) to represent non-administrative clinical interactions.
  • ChatGPT was prompted to generate responses with approximately the same word count as the original physician responses.
  • Participants were shown paired responses—five from providers, five from ChatGPT—and asked to identify the source.
  • Participants were financially incentivized to correctly classify each response source.
  • A 5-point Likert scale was used to measure trust in chatbots’ functions for patient communication.
  • A representative US sample of 430 adults aged 18+ was surveyed, with 53.2% women and average age 47.1.

Experimental results

Research questions

  • RQ1Can laypeople reliably distinguish between ChatGPT-generated and human-generated medical responses in patient-provider interactions?
  • RQ2How does user trust in chatbots for medical advice vary with the complexity of the health-related task?
  • RQ3To what extent does the perceived similarity between AI and human responses affect user confidence in AI for healthcare communication?
  • RQ4Are there differences in response classification accuracy across varying levels of clinical complexity?

Key findings

  • ChatGPT responses were correctly identified as AI-generated 65.5% of the time, indicating weak distinguishability from human responses.
  • Human provider responses were correctly classified 65.1% of the time, showing similar perceptual similarity to AI responses.
  • On average, participants rated their trust in chatbots’ medical functions at 3.4 on a 5-point Likert scale, indicating weakly positive but cautious trust.
  • Trust in chatbots decreased significantly as the health-related complexity of the query increased.
  • Classification accuracy varied widely by question, ranging from 49.0% to 85.7%, suggesting context-dependent detectability.
  • The results suggest that for low-risk, low-complexity health queries, users may accept chatbot responses as credible and trustworthy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.