[论文解读] Putting ChatGPT's Medical Advice to the (Turing) Test
本研究评估了ChatGPT在患者沟通中是否能令人信服地模拟人类医疗提供者,通过测试普通人能否区分AI生成与真实医生的回应。参与者平均正确识别回应来源的准确率为65.5%,表明ChatGPT的医疗建议与人类回应的可区分性较弱,且对低复杂度健康查询,用户对聊天机器人的信任度中等。
Objective: Assess the feasibility of using ChatGPT or a similar AI-based chatbot for patient-provider communication. Participants: A US representative sample of 430 study participants aged 18 and above. 53.2% of respondents analyzed were women; their average age was 47.1. Exposure: Ten representative non-administrative patient-provider interactions were extracted from the EHR. Patients' questions were placed in ChatGPT with a request for the chatbot to respond using approximately the same word count as the human provider's response. In the survey, each patient's question was followed by a provider- or ChatGPT-generated response. Participants were informed that five responses were provider-generated and five were chatbot-generated. Participants were asked, and incentivized financially, to correctly identify the response source. Participants were also asked about their trust in chatbots' functions in patient-provider communication, using a Likert scale of 1-5. Results: The correct classification of responses ranged between 49.0% to 85.7% for different questions. On average, chatbot responses were correctly identified 65.5% of the time, and provider responses were correctly distinguished 65.1% of the time. On average, responses toward patients' trust in chatbots' functions were weakly positive (mean Likert score: 3.4), with lower trust as the health-related complexity of the task in questions increased. Conclusions: ChatGPT responses to patient questions were weakly distinguishable from provider responses. Laypeople appear to trust the use of chatbots to answer lower risk health questions.
研究动机与目标
- 评估使用类似ChatGPT的AI聊天机器人进行患者-医生沟通的可行性。
- 评估普通人能否区分ChatGPT生成与人类生成的医疗回应。
- 检查用户对聊天机器人处理健康相关咨询的信任度。
- 探讨健康相关复杂性如何影响用户对AI医疗建议可信度的感知。
提出的方法
- 从电子健康记录(EHRs)中提取了十个真实患者问题,以代表非行政性的临床互动。
- 向ChatGPT提供提示,生成字数与原始医生回应大致相当的回应。
- 参与者被展示成对的回应——五条来自医生,五条来自ChatGPT——并被要求识别来源。
- 参与者因正确分类每条回应来源而获得经济激励。
- 使用5点李克特量表测量用户对聊天机器人在患者沟通中功能的信任度。
- 调查了430名18岁及以上的美国代表性样本,女性占53.2%,平均年龄47.1岁。
实验结果
研究问题
- RQ1普通人能否可靠地区分ChatGPT生成与人类生成的医疗回应?
- RQ2用户对聊天机器人在医疗建议方面的信任度如何随健康相关任务的复杂性而变化?
- RQ3AI与人类回应之间的感知相似性在多大程度上影响用户对AI在医疗沟通中可信度的信心?
- RQ4在不同临床复杂性水平下,回应分类的准确率是否存在差异?
主要发现
- ChatGPT的回应被正确识别为AI生成的准确率为65.5%,表明其与人类回应的可区分性较弱。
- 医生提供的回应被正确分类的准确率为65.1%,表明其与AI回应在感知上相似。
- 平均而言,参与者对聊天机器人医疗功能的信任度评分为3.4(5点李克特量表),表明信任度为弱正向但持谨慎态度。
- 随着查询的健康相关复杂性增加,用户对聊天机器人的信任度显著下降。
- 分类准确率因问题而异,范围在49.0%至85.7%之间,表明检测能力具有情境依赖性。
- 结果表明,对于低风险、低复杂度的健康查询,用户可能接受聊天机器人回应为可信且可靠。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。