[论文解读] The Potential and Pitfalls of using a Large Language Model such as ChatGPT or GPT-4 as a Clinical Assistant
本研究使用真实的电子健康记录(EHR)和模拟患者病例,评估了GPT-4和ChatGPT作为临床助手的表现。结果显示,GPT-4在使用思维链(chain-of-thought)和少样本提示(few-shot prompting)时,疾病分类的F1得分最高可达96%,但存在事实性错误、遗漏关键发现以及过度治疗风险,限制了其当前的临床实用性,尽管其具备良好的可扩展潜力。
Recent studies have demonstrated promising performance of ChatGPT and GPT-4 on several medical domain tasks. However, none have assessed its performance using a large-scale real-world electronic health record database, nor have evaluated its utility in providing clinical diagnostic assistance for patients across a full range of disease presentation. We performed two analyses using ChatGPT and GPT-4, one to identify patients with specific medical diagnoses using a real-world large electronic health record database and the other, in providing diagnostic assistance to healthcare workers in the prospective evaluation of hypothetical patients. Our results show that GPT-4 across disease classification tasks with chain of thought and few-shot prompting can achieve performance as high as 96% F1 scores. For patient assessment, GPT-4 can accurately diagnose three out of four times. However, there were mentions of factually incorrect statements, overlooking crucial medical findings, recommendations for unnecessary investigations and overtreatment. These issues coupled with privacy concerns, make these models currently inadequate for real world clinical use. However, limited data and time needed for prompt engineering in comparison to configuration of conventional machine learning workflows highlight their potential for scalability across healthcare applications.
研究动机与目标
- 评估GPT-4和ChatGPT在使用大规模真实世界电子健康记录(EHR)数据库时,识别特定医学诊断患者的表现。
- 评估这些大语言模型在为医疗专业人员提供假设性患者病例诊断辅助方面的实用性。
- 识别并分析在临床环境中部署大语言模型时出现的事实性错误、临床推理缺陷以及隐私风险等伦理问题。
- 比较提示工程与传统机器学习工作流在数据和时间需求上的差异,评估其可扩展潜力。
提出的方法
- 使用大规模真实世界EHR数据库进行回顾性分析,利用GPT-4结合思维链和少样本提示技术进行疾病分类的训练与评估。
- 设计一项前瞻性模拟研究,通过临床记录和症状信息提示GPT-4评估假设性患者病例。
- 采用少样本提示和思维链推理方法,以提升分类任务中模型的一致性和诊断准确性。
- 将模型输出与金标准临床诊断进行对比,计算F1分数并评估临床推理质量。
- 识别并分类模型响应中的错误,包括事实性错误、关键发现被遗漏,以及对不必要的检查或治疗的建议。
- 评估通过大语言模型交互可能导致EHR数据暴露所引发的隐私和安全问题。
实验结果
研究问题
- RQ1GPT-4是否能通过提示工程技术在真实世界EHR数据上实现高疾病分类诊断准确率?
- RQ2与标准临床推理相比,GPT-4在为假设性患者病例提供临床诊断辅助方面表现如何?
- RQ3作为临床助手使用GPT-4时,会涌现出哪些类型的事实性、临床性及伦理错误?
- RQ4提示工程在数据和时间需求方面与传统医疗领域机器学习工作流开发相比如何?
主要发现
- 当在真实EHR数据上使用思维链和少样本提示时,GPT-4在疾病分类任务中F1得分最高可达96%。
- 在前瞻性患者病例评估中,GPT-4准确诊断出四例中的三例,显示出强大的诊断潜力。
- 模型在大量响应中产生了事实性错误,包括对临床发现和指南的错误描述。
- GPT-4频繁遗漏患者摘要中的关键医学信息,如禁忌症或共病情况。
- 在多个病例中,模型建议了不必要的检查和治疗,表明存在过度治疗风险。
- 隐私问题仍是主要障碍,因为大语言模型交互可能导致敏感EHR数据暴露,限制其在现实世界中的部署。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。