[论文解读] Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries
本研究提出知识与少样本增强上下文学习(KFE)框架,通过整合外部临床知识和少样本范例,提升大语言模型(LLM)在非英语医学场景下的表现。KFE在中文国家医师资格考试(CNMLE-2022)上显著提升性能,使GPT-4得分达82.59/100,超过人类平均水平(68.70),并使参数量为13B的模型也能通过考试,证明该框架在低资源语言环境下的鲁棒性与可扩展性。
$ extbf{Objectives}$: Large Language Models (LLMs) such as ChatGPT and Med-PaLM have excelled in various medical question-answering tasks. However, these English-centric models encounter challenges in non-English clinical settings, primarily due to limited clinical knowledge in respective languages, a consequence of imbalanced training corpora. We systematically evaluate LLMs in the Chinese medical context and develop a novel in-context learning framework to enhance their performance. $ extbf{Materials and Methods}$: The latest China National Medical Licensing Examination (CNMLE-2022) served as the benchmark. We collected 53 medical books and 381,149 medical questions to construct the medical knowledge base and question bank. The proposed Knowledge and Few-shot Enhancement In-context Learning (KFE) framework leverages the in-context learning ability of LLMs to integrate diverse external clinical knowledge sources. We evaluated KFE with ChatGPT(GPT3.5), GPT4, Baichuan2(BC2)-7B, and BC2-13B in CNMLE-2022 and investigated the effectiveness of different pathways for incorporating LLMs with medical knowledge from 7 perspectives. $ extbf{Results}$: Directly applying ChatGPT failed to qualify for the CNMLE-2022 at a score of 51. Cooperated with the KFE, the LLMs with varying sizes yielded consistent and significant improvements. The ChatGPT's performance surged to 70.04 and GPT-4 achieved the highest score of 82.59. This surpasses the qualification threshold (60) and exceeds the average human score of 68.70. It also enabled a smaller BC2-13B to pass the examination, showcasing the great potential in low-resource settings. $ extbf{Conclusion}$: By synergizing medical knowledge through in-context learning, LLM can extend clinical insight beyond language barriers, significantly reducing language-related disparities of LLM applications and ensuring global benefit in healthcare.
研究动机与目标
- 为解决英语主导的LLM在非英语临床场景中因训练数据不平衡和多语言临床知识有限而导致的性能差距。
- 开发一种可扩展、无需训练的框架,通过结合外部医学知识和范例案例的上下文学习方式增强LLM。
- 评估多种知识整合路径在提升LLM于中文高阶医学考试中表现方面的有效性。
- 证明在整合精选临床知识与少样本示例后,较小的、低资源LLM可实现专家级表现。
- 通过在语言和资源受限环境中实现LLM的全球可及性,减少人工智能医疗应用中的语言相关差异。
提出的方法
- 从53本医学教科书和与CNMLE-2022对齐的381,149道题目题库中构建了全面的医学知识库。
- 设计KFE框架,检索相关临床知识和相似历史问题,形成上下文感知的提示,用于上下文学习。
- 将特定指令、检索到的知识、范例案例和目标问题整合到单一提示中,以引导LLM推理。
- 利用LLM的上下文学习能力,动态适应新输入,无需微调或重新训练。
- 在CNMLE-2022基准上,评估多种LLM(包括GPT-3.5、GPT-4、Baichuan2-7b和Baichuan2-13b)在KFE框架下的表现。
- 系统分析七种不同的知识整合路径,以识别最有效的性能增强配置。

实验结果
研究问题
- RQ1在非英语语言(如中文)的高阶医学执业考试中,外部临床知识与少样本范例能否显著提升LLM的性能?
- RQ2KFE框架如何在无需模型微调或重新训练的情况下提升LLM的临床推理能力?
- RQ3将多样化医学知识源整合到LLM中的最有效路径是什么,以实现性能的最大化提升?
- RQ4当整合精选知识与少样本范例后,较小的、低资源LLM是否能在医学考试中达到专家级表现?
- RQ5KFE框架在多大程度上可减少人工智能医疗应用中的语言与资源相关差异?
主要发现
- KFE框架使GPT-4在CNMLE-2022中取得82.59分,超过人类平均分68.70,且高于60分的及格线。
- ChatGPT(GPT-3.5)在KFE下得分从不及格的51分提升至70.04分,表现出持续且显著的性能提升。
- 尽管参数量较小且未针对医学任务进行充分预训练,Baichuan2-13B模型在集成KFE后通过了CNMLE-2022,凸显其在低资源环境中的潜力。
- 该框架在不同规模的LLM(从70亿到1.7万亿参数)中均表现出稳健性能,表明其具备广泛的可扩展性与泛化能力。
- 不同知识整合路径带来的性能提升各异,最优配置结合了知识检索、范例选择与结构化提示。
- 本研究证实,通过外部知识的上下文学习可有效突破语言边界,扩展临床洞察力,减少人工智能医疗中获取资源的不平等。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。