[论文解读] Large Language Models in Medical Term Classification and Unexpected Misalignment Between Response and Reasoning
本研究评估了 GPT-4、GPT-3.5、Falcon 和 LLaMA 2 等前沿大语言模型(LLMs)在使用 MIMIC-IV v2.2 数据集从临床出院摘要中分类轻度认知障碍(MCI)的表现。尽管 GPT-4 在推理和可解释性方面表现优异,但其推理过程与最终响应之间存在显著脱节,凸显了其在临床部署中可靠性方面存在关键缺陷,尽管其性能表现优异。
This study assesses the ability of state-of-the-art large language models (LLMs) including GPT-3.5, GPT-4, Falcon, and LLaMA 2 to identify patients with mild cognitive impairment (MCI) from discharge summaries and examines instances where the models' responses were misaligned with their reasoning. Utilizing the MIMIC-IV v2.2 database, we focused on a cohort aged 65 and older, verifying MCI diagnoses against ICD codes and expert evaluations. The data was partitioned into training, validation, and testing sets in a 7:2:1 ratio for model fine-tuning and evaluation, with an additional metastatic cancer dataset from MIMIC III used to further assess reasoning consistency. GPT-4 demonstrated superior interpretative capabilities, particularly in response to complex prompts, yet displayed notable response-reasoning inconsistencies. In contrast, open-source models like Falcon and LLaMA 2 achieved high accuracy but lacked explanatory reasoning, underscoring the necessity for further research to optimize both performance and interpretability. The study emphasizes the significance of prompt engineering and the need for further exploration into the unexpected reasoning-response misalignment observed in GPT-4. The results underscore the promise of incorporating LLMs into healthcare diagnostics, contingent upon methodological advancements to ensure accuracy and clinical coherence of AI-generated outputs, thereby improving the trustworthiness of LLMs for medical decision-making.
研究动机与目标
- 评估前沿大语言模型在从临床出院摘要中识别轻度认知障碍(MCI)方面的表现。
- 探究大语言模型生成的推理过程与其最终分类响应之间的一致性。
- 比较专有模型(如 GPT-4)与开源模型(如 Falcon、LLaMA 2)在可解释性和准确性方面的表现。
- 评估提示工程对模型行为及临床一致性的影响力。
- 识别在部署大语言模型用于可靠医疗诊断时面临的方法论挑战。
提出的方法
- 在 MIMIC-IV v2.2 数据集中,对 65 岁及以上、通过 ICD 编码和专家评审确诊为 MCI 的患者数据,采用 7:2:1 的训练/验证/测试集划分进行大语言模型微调。
- 使用复杂、多步骤的提示,以激发模型在分类前生成详细推理过程。
- 将模型输出与金标准 MCI 诊断进行对比,以检测推理与响应之间的不一致。
- 通过 MIMIC-III 数据集中转移性癌症数据集进行消融研究,以检验模型在不同领域中推理的一致性。
- 在准确性和可解释性方面,对比专有模型(如 GPT-4、GPT-3.5)与开源模型(如 Falcon、LLaMA 2)的性能表现。
- 采用定性与定量分析,衡量推理步骤与最终预测之间的一致性。
实验结果
研究问题
- RQ1大语言模型在从临床出院摘要中分类轻度认知障碍方面,准确度如何?
- RQ2大语言模型的推理过程与其最终分类响应之间的一致性程度如何?
- RQ3专有模型(如 GPT-4)与开源模型(如 Falcon 和 LLaMA 2)在准确性和可解释性方面表现如何比较?
- RQ4提示工程是否能提升大语言模型生成推理的一致性与临床一致性?
- RQ5推理一致性在不同医学疾病(如 MCI 和转移性癌症)之间是否具有可推广性?
主要发现
- GPT-4 展现出卓越的可解释能力,尤其在复杂提示场景下,但其推理与最终响应之间存在显著不一致。
- 开源模型如 Falcon 和 LLaMA 2 虽达到高分类准确率,但推理过程缺乏细节或一致性,限制了临床信任度。
- 尽管性能优异,GPT-4 的推理过程往往无法逻辑支持其最终预测,表明其在可靠性方面存在关键缺陷。
- 本研究发现 GPT-4 存在反复出现的推理-响应不一致现象,尤其在面对复杂或模糊的临床提示时更为明显。
- 在转移性癌症子集上的表现证实,推理不一致性在不同医学条件下持续存在,提示这是系统性问题。
- 研究结果强调,仅具备高准确率不足以支持临床部署;推理连贯性与可解释性是实现医疗领域可信人工智能的关键要素。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。