[Paper Review] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
MedXpertQA presents a highly challenging medical benchmark with text and multimodal subsets to evaluate expert-level medical reasoning, testing 18 models and introducing a reasoning-focused subset for o1-like models.
We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two subsets, Text for text evaluation and MM for multimodal evaluation. Notably, MM introduces expert-level exam questions with diverse images and rich clinical information, including patient records and examination results, setting it apart from traditional medical multimodal benchmarks with simple QA pairs generated from image captions. MedXpertQA applies rigorous filtering and augmentation to address the insufficient difficulty of existing benchmarks like MedQA, and incorporates specialty board questions to improve clinical relevance and comprehensiveness. We perform data synthesis to mitigate data leakage risk and conduct multiple rounds of expert reviews to ensure accuracy and reliability. We evaluate 18 leading models on \benchmark. Moreover, medicine is deeply connected to real-world decision-making, providing a rich and representative setting for assessing reasoning abilities beyond mathematics and code. To this end, we develop a reasoning-oriented subset to facilitate the assessment of o1-like models. Code and data are available at: https://github.com/TsinghuaC3I/MedXpertQA
Motivation & Objective
- Assess expert-level medical knowledge and reasoning across diverse specialties and body systems.
- Provide both text-only and multimodal evaluation to reflect real-world clinical tasks.
- Increase difficulty and realism over existing benchmarks through filtering, augmentation, and expert review.
- Enable analysis of advanced medical reasoning capabilities beyond basic knowledge or perception.
Proposed method
- Curate a large question bank from USMLE, COMLEX-USA, specialty board exams, and image-rich sources like NEJM Image Challenges.
- Apply hierarchical difficulty-based and expert filtering to select challenging questions using Brier scores and expert votes.
- Augment questions and options with LLMs to enhance diversity and difficulty while mitigating data leakage.
- Introduce data synthesis and multiple rounds of expert review to ensure accuracy and validity.
- Annotate questions with core medical tasks (Diagnosis, Treatment Planning, Basic Medicine) and fine-grained subtasks using GPT-4o prompts.
- Evaluate 18 large models (LMMs and LLMs) on zero-shot CoT prompting, with dedicated Reasoning and Understanding subsets for specialized analysis.
Experimental results
Research questions
- RQ1How capable are state-of-the-art LMMs/LLMs at expert-level medical reasoning across diverse specialties and body systems?
- RQ2Can a benchmark like MedXpertQA reliably assess reasoning beyond factual medical knowledge, including multimodal tasks?
- RQ3What is the performance gap between current models and expert human baselines on reasoning-intensive medical questions?
- RQ4How does inference-time scaling affect medical reasoning performance on challenging tasks?
- RQ5Can a Reasoning-focused subset effectively differentiate o1-like reasoning models in medicine?
Key findings
- MedXpertQA is highly challenging, with current models showing limited performance on complex medical reasoning tasks.
- GPT-4o generally performs best among vanilla LMMs on MedXpertQA, with GPT-4o-mini and other models following; Qwen-series and open models vary in multimodal settings.
- DeepSeek-R1 shows strong reasoning performance, especially on the Reasoning subset, highlighting the need for specialized evaluation of medical reasoning abilities.
- The Reasoning subset is reliably more challenging than Understanding, indicating effective differentiation of reasoning capabilities in o1-like models.
- Data augmentation and multi-round expert review reduce leakage risk and improve dataset quality without compromising core clinical content.
- MedXpertQA MM demonstrates higher complexity and richer imagery, while MedXpertQA Text emphasizes specialty-driven evaluation across 11 body systems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.