Skip to main content
QUICK REVIEW

[论文解读] People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy

Shruthi Shekar, Pat Pataranutaporn|arXiv (Cornell University)|Aug 11, 2024
Artificial Intelligence in Healthcare and Education被引用 11
一句话总结

该研究表明非专业人士无法区分 AI 生成的医疗回复与医生的回复,往往信任 AI 与医生同样有效或更可信,特别是当 AI 标注为高准确性时,引发安全担忧。

ABSTRACT

This paper presents a comprehensive analysis of how AI-generated medical responses are perceived and evaluated by non-experts. A total of 300 participants gave evaluations for medical responses that were either written by a medical doctor on an online healthcare platform, or generated by a large language model and labeled by physicians as having high or low accuracy. Results showed that participants could not effectively distinguish between AI-generated and Doctors' responses and demonstrated a preference for AI-generated responses, rating High Accuracy AI-generated responses as significantly more valid, trustworthy, and complete/satisfactory. Low Accuracy AI-generated responses on average performed very similar to Doctors' responses, if not more. Participants not only found these low-accuracy AI-generated responses to be valid, trustworthy, and complete/satisfactory but also indicated a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result of the response provided. This problematic reaction was comparable if not more to the reaction they displayed towards doctors' responses. This increased trust placed on inaccurate or inappropriate AI-generated medical advice can lead to misdiagnosis and harmful consequences for individuals seeking help. Further, participants were more trusting of High Accuracy AI-generated responses when told they were given by a doctor and experts rated AI-generated responses significantly higher when the source of the response was unknown. Both experts and non-experts exhibited bias, finding AI-generated responses to be more thorough and accurate than Doctors' responses but still valuing the involvement of a Doctor in the delivery of their medical advice. Ensuring AI systems are implemented with medical professionals should be the future of using AI for the delivery of medical advice.

研究动机与目标

  • 评估普通人是否能够区分 AI 生成的医疗回答与医生提供的回答。
  • 评估对 AI 与医生回答在有效性、可信度、完整性和用户意向方面的感知。
  • 考察对答案来源的知晓程度如何影响感知与信任。
  • 探讨当来源揭示或未知时,专家评估者是否会出现偏见。

提出的方法

  • 收集 150 条 AI 生成的 HealthTap 问题回答,并让四位医生对准确性进行打分(是、可能、否),以将 AI 输出分为高准确性与低准确性。
  • 创建一个包含 30 条医生回答、30 条高准确性 AI 回答、30 条低准确性 AI 回答的数据集,针对 100 名在线参与者。
  • 实验 1:参与者判断理解与来源,并对来源的自信度进行比较;AI 与医生对比。
  • 实验 2:参与者在不知道确切来源的情况下评估 AI 与医生;测量有效性、信任、完整性和行为意向。
  • 实验 3:随机标签实验,测试来源描述(医生、AI、由 AI 辅助的医生)对评估的偏见。
  • 在盲/非盲条件下的额外医生评估,以评估在披露来源时专家对 AI 的偏见。
Figure 1: Visual summary of the dataset construction and pipeline of experiments discussed in this paper.
Figure 1: Visual summary of the dataset construction and pipeline of experiments discussed in this paper.

实验结果

研究问题

  • RQ1参与者是否能够区分 AI 生成的医疗回答与医生提供的回答?
  • RQ2参与者对 AI 与医生回答在有效性、信任、完整性和行为意向方面的评价如何?
  • RQ3知晓回答来源(医生 vs. AI)是否会影响感知与信任?
  • RQ4专家在评估有无来源披露的 AI 输出时是否存在偏见?

主要发现

  • 参与者无法可靠地区分 AI 生成的回答与医生的回答,在不同类型之间的来源识别准确性约为 50%。
  • 在实验 2 中,高准确性 AI 生成的回答被评为更具有效性、可信度和完整性/满意度,显著高于医生回答。
  • 低准确性 AI 生成的回答在有效性、信任和完整性方面的表现与医生的回答相似;在某些情况下甚至超过医生的表现。
  • 当来源未知时,参与者通常信任 AI 生成的回答,但当标注为来自医生时,对高准确性 AI 的信任度进一步提高。
  • 专家在来源未知时对 AI 生成的回答给出更高的评分,但在获知回答来自 AI 时,其评分显著下降。
  • 研究强调了错误信息的风险:普通用户若未经临床医生监督,可能会遵循有害的 AI 建议。
Figure 2: Example Medical Questions by Category: Comparing Doctors’ and AI-Generated Responses
Figure 2: Example Medical Questions by Category: Comparing Doctors’ and AI-Generated Responses

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。