[论文解读] A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology
本文将开源大型语言模型与GPT-4和Claude 2在肾脏病学多项选择题上的表现进行比较,结果显示GPT-4和Claude 2在大幅领先开源模型。
In recent years, there have been significant breakthroughs in the field of natural language processing, particularly with the development of large language models (LLMs). These LLMs have showcased remarkable capabilities on various benchmarks. In the healthcare field, the exact role LLMs and other future AI models will play remains unclear. There is a potential for these models in the future to be used as part of adaptive physician training, medical co-pilot applications, and digital patient interaction scenarios. The ability of AI models to participate in medical training and patient care will depend in part on their mastery of the knowledge content of specific medical fields. This study investigated the medical knowledge capability of LLMs, specifically in the context of internal medicine subspecialty multiple-choice test-taking ability. We compared the performance of several open-source LLMs (Koala 7B, Falcon 7B, Stable-Vicuna 13B, and Orca Mini 13B), to GPT-4 and Claude 2 on multiple-choice questions in the field of Nephrology. Nephrology was chosen as an example of a particularly conceptually complex subspecialty field within internal medicine. The study was conducted to evaluate the ability of LLM models to provide correct answers to nephSAP (Nephrology Self-Assessment Program) multiple-choice questions. The overall success of open-sourced LLMs in answering the 858 nephSAP multiple-choice questions correctly was 17.1% - 25.5%. In contrast, Claude 2 answered 54.4% of the questions correctly, whereas GPT-4 achieved a score of 73.3%. We show that current widely used open-sourced LLMs do poorly in their ability for zero-shot reasoning when compared to GPT-4 and Claude 2. The findings of this study potentially have significant implications for the future of subspecialty medical training and patient care.
研究动机与目标
- 评估 open-source LLMs 在 nephSAP MCQs 上的肾脏病学知识能力。
- 在肾脏病学数据集上对比 open-source LLMs 与 GPT-4 与 Claude 2 的零-shot 表现。
- 分析模型解释的质量及其与地面真相的语义对齐。
提出的方法
- 对 Koala 7B、Falcon 7B、Stable-Vicuna 13B 和 Orca Mini 13B 在 858 道 nephSAP MCQs 上与 GPT-4 和 Claude 2 进行对比评估。
- 通过将 Context、Question 和 Choices 拼接成输入提示,用于开源 LLM 的前向传播。
- 使用基于正则表达式的提取对自动输出进行解析,以确定预测答案并与地面真相进行比较。
- 将准确率衡量为原始正确回答数,并计算各肾脏病学子主题的分主题表现。
- 使用 BLEU、WER 与余弦相似度来评估对地面真相解释的质量。
实验结果
研究问题
- RQ1开源 LLMs 在 nephSAP 肾脏病学多项选择题上的表现与 GPT-4 及 Claude 2 相比如何?
- RQ2开源 LLM 在肾脏病学中的按主题的优势和弱点是什么?
- RQ3开源 LLMs 与地面真相相比,在解释质量方面的表现如何?
- RQ4可能影响开源 LLM 与 GPT-4/Claude 2 之间性能差距的因素有哪些?
主要发现
- GPT-4 在正确答案数量上达到 629 道(73.3%),Claude 2 为 467 道(54.4%),而开源 LLMs 的正确率范围为 17.1% 至 25.5%。
- 在开源模型中,Vicuna 的得分最高,为 219 道(25.5%),Koala 跟着为 204 道(23.8%)。
- Falcon 获得 155 道正确(18.1%),Orca-Mini 147 道正确(17.1%),Koala 204 道正确(23.8%)。
- 在问题结构下,随机猜测的正确率为 23.8%,因此在开源模型中只有 Koala 稍微超出随机期望。
- 开源模型的解释显示较低的 BLEU 分数(例如 Vicuna 9%、Falcon 8%、Orca-Mini 5%、Koala 5%),且各模型的余弦相似度分数也普遍不理想。
- GPT-4 与 Claude 2 在所有肾脏病学主题上均优于开源模型,且 GPT-4 在大多数主题上接近人类表现。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。