[Paper Review] Putting ChatGPT's Medical Advice to the (Turing) Test
This study evaluates whether ChatGPT can convincingly emulate human medical providers in patient communication by testing if laypeople can distinguish between AI-generated and real physician responses. Participants correctly identified response sources at 65.5% accuracy on average, indicating ChatGPT’s medical advice is weakly distinguishable from human responses, with moderate trust in chatbots for low-complexity health queries.
Objective: Assess the feasibility of using ChatGPT or a similar AI-based chatbot for patient-provider communication. Participants: A US representative sample of 430 study participants aged 18 and above. 53.2% of respondents analyzed were women; their average age was 47.1. Exposure: Ten representative non-administrative patient-provider interactions were extracted from the EHR. Patients' questions were placed in ChatGPT with a request for the chatbot to respond using approximately the same word count as the human provider's response. In the survey, each patient's question was followed by a provider- or ChatGPT-generated response. Participants were informed that five responses were provider-generated and five were chatbot-generated. Participants were asked, and incentivized financially, to correctly identify the response source. Participants were also asked about their trust in chatbots' functions in patient-provider communication, using a Likert scale of 1-5. Results: The correct classification of responses ranged between 49.0% to 85.7% for different questions. On average, chatbot responses were correctly identified 65.5% of the time, and provider responses were correctly distinguished 65.1% of the time. On average, responses toward patients' trust in chatbots' functions were weakly positive (mean Likert score: 3.4), with lower trust as the health-related complexity of the task in questions increased. Conclusions: ChatGPT responses to patient questions were weakly distinguishable from provider responses. Laypeople appear to trust the use of chatbots to answer lower risk health questions.
Motivation & Objective
- To evaluate the feasibility of using AI chatbots like ChatGPT for patient-provider communication.
- To assess whether laypeople can distinguish between ChatGPT-generated and human-generated medical responses.
- To examine user trust in chatbots for handling health-related inquiries.
- To explore how health-related complexity affects perceived trustworthiness of AI medical advice.
Proposed method
- Ten real patient questions were extracted from electronic health records (EHRs) to represent non-administrative clinical interactions.
- ChatGPT was prompted to generate responses with approximately the same word count as the original physician responses.
- Participants were shown paired responses—five from providers, five from ChatGPT—and asked to identify the source.
- Participants were financially incentivized to correctly classify each response source.
- A 5-point Likert scale was used to measure trust in chatbots’ functions for patient communication.
- A representative US sample of 430 adults aged 18+ was surveyed, with 53.2% women and average age 47.1.
Experimental results
Research questions
- RQ1Can laypeople reliably distinguish between ChatGPT-generated and human-generated medical responses in patient-provider interactions?
- RQ2How does user trust in chatbots for medical advice vary with the complexity of the health-related task?
- RQ3To what extent does the perceived similarity between AI and human responses affect user confidence in AI for healthcare communication?
- RQ4Are there differences in response classification accuracy across varying levels of clinical complexity?
Key findings
- ChatGPT responses were correctly identified as AI-generated 65.5% of the time, indicating weak distinguishability from human responses.
- Human provider responses were correctly classified 65.1% of the time, showing similar perceptual similarity to AI responses.
- On average, participants rated their trust in chatbots’ medical functions at 3.4 on a 5-point Likert scale, indicating weakly positive but cautious trust.
- Trust in chatbots decreased significantly as the health-related complexity of the query increased.
- Classification accuracy varied widely by question, ranging from 49.0% to 85.7%, suggesting context-dependent detectability.
- The results suggest that for low-risk, low-complexity health queries, users may accept chatbot responses as credible and trustworthy.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.