[论文解读] Medical Misinformation in AI-Assisted Self-Diagnosis: Development of a Method (EvalPrompt) for Analyzing Large Language Models
本文提出 EvalPrompt,一种新颖方法,通过在无答案选项的开放式医学问题中测试大型语言模型(LLMs)如 ChatGPT,来评估其在真实自我诊断场景中的表现。结果表明,LLMs 的表现显著低于以往假设,对错误建议表现出过度自信,且未能包含免责声明或表明不确定性——这在实际应用中加剧了医学误诊信息传播的风险。
Rapid integration of large language models (LLMs) in health care is sparking global discussion about their potential to revolutionize health care quality and accessibility. At a time when improving health care quality and access remains a critical concern for countries worldwide, the ability of these models to pass medical examinations is often cited as a reason to use them for medical training and diagnosis. However, the impact of their inevitable use as a self-diagnostic tool and their role in spreading healthcare misinformation has not been evaluated. This study aims to assess the effectiveness of LLMs, particularly ChatGPT, from the perspective of an individual self-diagnosing to better understand the clarity, correctness, and robustness of the models. We propose the comprehensive testing methodology evaluation of LLM prompts (EvalPrompt). This evaluation methodology uses multiple-choice medical licensing examination questions to evaluate LLM responses. We use open-ended questions to mimic real-world self-diagnosis use cases, and perform sentence dropout to mimic realistic self-diagnosis with missing information. Human evaluators then assess the responses returned by ChatGPT for both experiments for clarity, correctness, and robustness. The results highlight the modest capabilities of LLMs, as their responses are often unclear and inaccurate. As a result, medical advice by LLMs should be cautiously approached. However, evidence suggests that LLMs are steadily improving and could potentially play a role in healthcare systems in the future. To address the issue of medical misinformation, there is a pressing need for the development of a comprehensive self-diagnosis dataset. This dataset could enhance the reliability of LLMs in medical applications by featuring more realistic prompt styles with minimal information across a broader range of medical fields.
研究动机与目标
- 批判性评估类似 ChatGPT 的 LLM 在自我诊断中的实际表现,超越多项选择题考试基准。
- 研究非专业人士在开放式症状查询中使用 LLM 时,其如何导致医学误诊信息传播。
- 开发一种可重复的、基于人工评估者的评估框架,以捕捉超越简单正确性的细微响应质量。
- 检验 LLM 是否能识别不确定性或验证自身响应,特别是在高风险医学情境中。
- 为医疗应用中 LLM 的评估提供方法论蓝图,优先考虑实际可用性而非考试表现。
提出的方法
- 将部分单答案 USMLE Step 1 问题改编为开放式提示,以模拟真实用户自我诊断场景。
- 使用非专业人工评估者,按细致评分标准对响应进行评级:正确、部分正确、错误或模糊。
- 通过句子删除(消融)进行敏感性分析,测试 LLM 响应在缺失信息情况下的鲁棒性。
- 要求 ChatGPT 对其自身答案进行自我评估,以衡量其识别错误或不确定性的能力。
- 利用汇总的人工评级量化 LLM 输出在性能下降和行为异常方面的表现。
- 建立一种可重复的、考虑编码者一致性的方法论,用于评估 LLM 在开放式医学推理任务中的表现。
实验结果
研究问题
- RQ1当不提供答案选项时,类似 ChatGPT 的 LLM 在医学自我诊断任务中的表现相较于多项选择题设置有何变化?
- RQ2LLM 在其医学响应中未能包含免责声明或表明不确定性的情况有多严重,即使其建议是错误的?
- RQ3信息缺失(关键临床细节被移除)如何影响 LLM 生成的医学建议的感知正确性?
- RQ4LLM 是否能准确自我评估其自身响应,特别是识别出错误或过度自信的回答?
- RQ5在医学推理任务中,LLM 响应中会浮现哪些行为模式——例如过度自信或不必要的谨慎?
主要发现
- 当不提供答案选项时,LLM 表现显著下降,表明多项选择题基准高估了其在真实世界中的诊断能力。
- ChatGPT 经常未能包含免责声明或表明不确定性,即使其建议是错误或可能有害的医学行为。
- 该模型在错误建议中表现出过度自信,增加了在自我诊断场景中传播医学误诊信息的风险。
- 当正确医学答案中包含额外但非必要的临床建议时,ChatGPT 常常将其归类为“部分正确”,表明其缺乏临床直觉。
- 在自我评估时,ChatGPT 对自身错误缺乏意识,表明其无法可靠地充当事实核查工具。
- 即使轻微的信息缺失(例如删除单一句子)也会显著改变响应的感知正确性,凸显其在现实世界数据变化下的推理脆弱性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。