Skip to main content
QUICK REVIEW

[论文解读] Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery

Debadutta Dash, Rahul Thapa|arXiv (Cornell University)|Apr 26, 2023
Artificial Intelligence in Healthcare and Education被引用 21
一句话总结

本研究评估 GPT-3.5 和 GPT-4 作为信息学咨询辅助工具,帮助医生在医疗保健中的提问,结果显示与专家报告的一致性有限,未出现压倒性的有害迹象;强调提示工程和模型定制的必要性。

ABSTRACT

Despite growing interest in using large language models (LLMs) in healthcare, current explorations do not assess the real-world utility and safety of LLMs in clinical settings. Our objective was to determine whether two LLMs can serve information needs submitted by physicians as questions to an informatics consultation service in a safe and concordant manner. Sixty six questions from an informatics consult service were submitted to GPT-3.5 and GPT-4 via simple prompts. 12 physicians assessed the LLM responses' possibility of patient harm and concordance with existing reports from an informatics consultation service. Physician assessments were summarized based on majority vote. For no questions did a majority of physicians deem either LLM response as harmful. For GPT-3.5, responses to 8 questions were concordant with the informatics consult report, 20 discordant, and 9 were unable to be assessed. There were 29 responses with no majority on "Agree", "Disagree", and "Unable to assess". For GPT-4, responses to 13 questions were concordant, 15 discordant, and 3 were unable to be assessed. There were 35 responses with no majority. Responses from both LLMs were largely devoid of overt harm, but less than 20% of the responses agreed with an answer from an informatics consultation service, responses contained hallucinated references, and physicians were divided on what constitutes harm. These results suggest that while general purpose LLMs are able to provide safe and credible responses, they often do not meet the specific information need of a given question. A definitive evaluation of the usefulness of LLMs in healthcare settings will likely require additional research on prompt engineering, calibration, and custom-tailoring of general purpose models.

研究动机与目标

  • 评估两种大型语言模型(GPT-3.5 和 GPT-4)是否能够安全回答医生提交的信息学问题。
  • 评估语言模型回答与已建立的信息学咨询报告之间的一致性。
  • 识别现实世界临床查询中的安全隐患,包括潜在的患者伤害和幻觉性信息。

提出的方法

  • 将来自信息学咨询服务的66个医生问题以简单提示提交给 GPT-3.5 和 GPT-4。
  • 请12位医生评估语言模型回答是否构成对患者的伤害,以及是否与信息学咨询报告一致。
  • 以多数投票方式总结医生评估,以确定安全性和一致性。
  • 报告每个模型的符合、一致性不足和无法评估的计数。

实验结果

研究问题

  • RQ1GPT-3.5 和 GPT-4 是否能够为现实世界的医生信息需求在医疗服务中提供安全的回答?
  • RQ2LLM 的回答在多大程度上与已建立的信息学咨询报告一致?
  • RQ3在临床查询的 LLM 输出中观察到的有害、幻觉性或不一致的模式是什么?

主要发现

  • 没有医生在任何问题上认定任意 LLM 的回答构成有害。
  • GPT-3.5:8 一致,20 不一致,9 无法评估;在 Agree/Disagree/Unable 三项上无多数的共有 29 条。
  • GPT-4:13 一致,15 不一致,3 无法评估;在 Agree/Disagree/Unable 三项上无多数的共有 35 条。
  • 两者的回答大多未显现明确的有害性,但包含幻觉式引用,且经常与信息学咨询报告不一致。
  • 不到 20% 的回答与信息学咨询服务的答案一致。
  • 表明通用 LLMs 在安全方面是可行的,但若不进行进一步的提示工程和定制,实用性并不可靠。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。