[论文解读] People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy
该研究表明非专业人士无法区分 AI 生成的医疗回复与医生的回复,往往信任 AI 与医生同样有效或更可信,特别是当 AI 标注为高准确性时,引发安全担忧。
This paper presents a comprehensive analysis of how AI-generated medical responses are perceived and evaluated by non-experts. A total of 300 participants gave evaluations for medical responses that were either written by a medical doctor on an online healthcare platform, or generated by a large language model and labeled by physicians as having high or low accuracy. Results showed that participants could not effectively distinguish between AI-generated and Doctors' responses and demonstrated a preference for AI-generated responses, rating High Accuracy AI-generated responses as significantly more valid, trustworthy, and complete/satisfactory. Low Accuracy AI-generated responses on average performed very similar to Doctors' responses, if not more. Participants not only found these low-accuracy AI-generated responses to be valid, trustworthy, and complete/satisfactory but also indicated a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result of the response provided. This problematic reaction was comparable if not more to the reaction they displayed towards doctors' responses. This increased trust placed on inaccurate or inappropriate AI-generated medical advice can lead to misdiagnosis and harmful consequences for individuals seeking help. Further, participants were more trusting of High Accuracy AI-generated responses when told they were given by a doctor and experts rated AI-generated responses significantly higher when the source of the response was unknown. Both experts and non-experts exhibited bias, finding AI-generated responses to be more thorough and accurate than Doctors' responses but still valuing the involvement of a Doctor in the delivery of their medical advice. Ensuring AI systems are implemented with medical professionals should be the future of using AI for the delivery of medical advice.
研究动机与目标
- 评估普通人是否能够区分 AI 生成的医疗回答与医生提供的回答。
- 评估对 AI 与医生回答在有效性、可信度、完整性和用户意向方面的感知。
- 考察对答案来源的知晓程度如何影响感知与信任。
- 探讨当来源揭示或未知时,专家评估者是否会出现偏见。
提出的方法
- 收集 150 条 AI 生成的 HealthTap 问题回答,并让四位医生对准确性进行打分(是、可能、否),以将 AI 输出分为高准确性与低准确性。
- 创建一个包含 30 条医生回答、30 条高准确性 AI 回答、30 条低准确性 AI 回答的数据集,针对 100 名在线参与者。
- 实验 1:参与者判断理解与来源,并对来源的自信度进行比较;AI 与医生对比。
- 实验 2:参与者在不知道确切来源的情况下评估 AI 与医生;测量有效性、信任、完整性和行为意向。
- 实验 3:随机标签实验,测试来源描述(医生、AI、由 AI 辅助的医生)对评估的偏见。
- 在盲/非盲条件下的额外医生评估,以评估在披露来源时专家对 AI 的偏见。

实验结果
研究问题
- RQ1参与者是否能够区分 AI 生成的医疗回答与医生提供的回答?
- RQ2参与者对 AI 与医生回答在有效性、信任、完整性和行为意向方面的评价如何?
- RQ3知晓回答来源(医生 vs. AI)是否会影响感知与信任?
- RQ4专家在评估有无来源披露的 AI 输出时是否存在偏见?
主要发现
- 参与者无法可靠地区分 AI 生成的回答与医生的回答,在不同类型之间的来源识别准确性约为 50%。
- 在实验 2 中,高准确性 AI 生成的回答被评为更具有效性、可信度和完整性/满意度,显著高于医生回答。
- 低准确性 AI 生成的回答在有效性、信任和完整性方面的表现与医生的回答相似;在某些情况下甚至超过医生的表现。
- 当来源未知时,参与者通常信任 AI 生成的回答,但当标注为来自医生时,对高准确性 AI 的信任度进一步提高。
- 专家在来源未知时对 AI 生成的回答给出更高的评分,但在获知回答来自 AI 时,其评分显著下降。
- 研究强调了错误信息的风险:普通用户若未经临床医生监督,可能会遵循有害的 AI 建议。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。