Skip to main content
QUICK REVIEW

[论文解读] MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Yan Cai, Linlin Wang|arXiv (Cornell University)|Dec 20, 2023
Topic Modeling被引用 4
一句话总结

MedBench 是一个大规模、多源的中文医学问答基准,包含来自官方考试和真实临床病例的 40,041 个中文医学问题,旨在评估医学大语言模型在完整中文医学训练体系下的知识与推理能力。结果表明,中文医学大语言模型表现显著不足,而一些通用领域模型(如 ChatGPT)则展现出强大的医学知识能力,凸显了在临床推理与诊断准确性方面的显著差距。

ABSTRACT

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the Chinese medical domain, comprising 40,041 questions sourced from authentic examination exercises and medical reports of diverse branches of medicine. In particular, this benchmark is composed of four key components: the Chinese Medical Licensing Examination, the Resident Standardization Training Examination, the Doctor In-Charge Qualification Examination, and real-world clinic cases encompassing examinations, diagnoses, and treatments. MedBench replicates the educational progression and clinical practice experiences of doctors in Mainland China, thereby establishing itself as a credible benchmark for assessing the mastery of knowledge and reasoning abilities in medical language learning models. We perform extensive experiments and conduct an in-depth analysis from diverse perspectives, which culminate in the following findings: (1) Chinese medical LLMs underperform on this benchmark, highlighting the need for significant advances in clinical knowledge and diagnostic precision. (2) Several general-domain LLMs surprisingly possess considerable medical knowledge. These findings elucidate both the capabilities and limitations of LLMs within the context of MedBench, with the ultimate goal of aiding the medical research community.

研究动机与目标

  • 为解决缺乏标准化、大规模评估基准的问题,以真实反映中国大陆医学教育与临床实践标准,针对中文医学大语言模型进行评估。
  • 通过纳入多样且真实的来源(如中国医师资格考试、住院医师规范化培训考试、主任医师资格考试等),克服现有基准的局限性。
  • 整合电子病历中的真实临床病例,以评估超越教科书知识的实用诊断推理能力。
  • 采用项目反应理论(Item Response Theory)构建具有心理测量学验证的基准,确保评估的可靠性与可扩展性。
  • 识别中文专用模型与通用领域大语言模型在临床知识回忆、推理与解释生成方面的优劣势。

提出的方法

  • 从四个真实来源构建 MedBench:中国医师资格考试、住院医师规范化培训考试、主任医师资格考试以及真实临床病例,共包含 40,041 个问题。
  • 将基准按四个医学分支(内科、外科、妇产科、儿科)组织,并采用从 Level 1 到 Level 9 的分层难度分级体系。
  • 整合多种推理类型:单跳推理、陈述识别与多跳推理,以评估逻辑与诊断推理能力。
  • 收集模型生成的答案解释,以评估推理过程的一致性与正确性。
  • 应用项目反应理论(IRT)建模题目难度与模型表现,实现更可靠、可扩展的评估。
  • 开展人工评估以验证模型输出,评估对话质量与临床推理准确性。
Figure 1: Comparison of procedures in different countries.
Figure 1: Comparison of procedures in different countries.

实验结果

研究问题

  • RQ1中文医学大语言模型在全面反映中国大陆医生教育与临床发展路径的基准上表现如何?
  • RQ2通用领域大语言模型(如 ChatGPT)在中文语境下是否具备可迁移的医学知识,相较于领域专用模型表现如何?
  • RQ3大语言模型在生成答案解释时的逻辑性如何?解释质量是否与答案正确性相关?
  • RQ4当前中文医学大语言模型在诊断推理与临床知识回忆方面存在哪些关键局限?
  • RQ5心理测量学方法(如项目反应理论)能否提升医学大语言模型评估的可靠性与可扩展性?

主要发现

  • 中文医学大语言模型在 MedBench 上表现显著不足,表明其在临床知识与诊断精度方面仍有巨大提升空间。
  • 多个通用领域大语言模型(包括 ChatGPT)展现出强大的医学知识与推理能力,表明通用模型向医学领域存在潜在的知识迁移能力。
  • 人工评估显示,尽管 ChatGPT 具有丰富的临床知识,但当前中文医学大语言模型在高质量对话能力与事实准确性方面仍显不足。
  • 当模型给出错误答案时,常伴随不合逻辑的解释,表明其存在知识空白或推理机制缺陷。
  • ChatGPT 在复杂病例分析题上的表现欠佳,凸显其在深层医学理解与临床推理方面存在关键缺陷。
  • 将项目反应理论整合至基准中,可实现更可靠、可扩展的模型评估,为未来评估方法的优化提供支持。
Figure 2: Examples of prompts and corresponding answers. The left side shows the prompt and responses of LLMs on an example question from an exam. The right side is an example of a real-world case.
Figure 2: Examples of prompts and corresponding answers. The left side shows the prompt and responses of LLMs on an example question from an exam. The right side is an example of a real-world case.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。