Skip to main content
QUICK REVIEW

[论文解读] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Yuxin Zuo, Shang Qu|ArXiv.org|Jan 30, 2025
Biomedical Text Mining and Ontologies被引用 3
一句话总结

MedXpertQA 提供了一个极具挑战性的医疗基准,包含文本与多模态子集,用于评估专家级医疗推理,测试 18 个模型并引入一个面向 o1-like 模型的推理聚焦子集。

ABSTRACT

We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two subsets, Text for text evaluation and MM for multimodal evaluation. Notably, MM introduces expert-level exam questions with diverse images and rich clinical information, including patient records and examination results, setting it apart from traditional medical multimodal benchmarks with simple QA pairs generated from image captions. MedXpertQA applies rigorous filtering and augmentation to address the insufficient difficulty of existing benchmarks like MedQA, and incorporates specialty board questions to improve clinical relevance and comprehensiveness. We perform data synthesis to mitigate data leakage risk and conduct multiple rounds of expert reviews to ensure accuracy and reliability. We evaluate 18 leading models on \benchmark. Moreover, medicine is deeply connected to real-world decision-making, providing a rich and representative setting for assessing reasoning abilities beyond mathematics and code. To this end, we develop a reasoning-oriented subset to facilitate the assessment of o1-like models. Code and data are available at: https://github.com/TsinghuaC3I/MedXpertQA

研究动机与目标

  • 在不同专业和体系统中评估专家级医学知识与推理能力。
  • 提供文本-only 与多模态评估,以反映真实世界临床任务。
  • 通过筛选、增量扩充和专家评审提高难度和现实性,超越现有基准。
  • 使分析超出基础知识或感知能力的高级医学推理能力。

提出的方法

  • 从USMLE、COMLEX-USA、专科板考试和如NEJM Image Challenges等富含图像来源中整理大量题库。
  • 应用基于层级难度的筛选和专家筛选,利用 Brier 分数和专家投票选出具有挑战性的题目。
  • 使用 LLM 对题目和选项进行增量扩充,以提高多样性和难度,同时降低数据泄露的风险。
  • 引入数据综合和多轮专家评审以确保准确性和有效性。
  • 使用 GPT-4o 提示对题目进行核心医学任务(诊断、治疗计划、基础医学)及细粒度子任务的标注。
  • 在零-shot 连锁推理提示下评估 18 个大型模型(LMMs 与 LLMs),并为专门分析设立专门的 Reasoning 与 Understanding 子集。

实验结果

研究问题

  • RQ1最先进的 LMMs/LLMs 在跨越多专业和体系统的专家级医疗推理方面有多大能力?
  • RQ2像 MedXpertQA 这样的基准能否可靠评估超越事实性医学知识的推理能力,包括多模态任务?
  • RQ3当前模型在推理密集型医学问题上的表现与专家人类基线之间的差距有多大?
  • RQ4推理阶段的扩展如何影响对具有挑战性任务的医学推理表现?
  • RQ5一个以推理为重点的子集是否能有效区分医学领域的 o1-like 推理模型?

主要发现

  • MedXpertQA 极具挑战性,当前模型在复杂医学推理任务上的表现有限。
  • 在 MedXpertQA 上,GPT-4o 通常在普通 LMMs 中表现最佳,其次是 GPT-4o-mini 及其他模型;Qwen 系列与开源模型在多模态设置中表现各异。
  • DeepSeek-R1 显示出较强的推理能力,尤其在 Reasoning 子集上,凸显了对医学推理能力进行专门评估的必要性。
  • Reasoning 子集在难度上明显高于 Understanding,能够有效区分 o1-like 模型的推理能力。
  • 数据增强与多轮专家评审降低了数据泄露风险并提升数据集质量,同时不损害核心临床内容。
  • MedXpertQA MM 展示了更高的复杂性与更丰富的图像,而 MedXpertQA Text 则强调跨 11 个体系统的专业驱动评估。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。