Skip to main content
QUICK REVIEW

[Paper Review] Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in Medicine

Maxime Griot, Jean Vanderdonckt|arXiv (Cornell University)|Jun 4, 2024
Expert finding and Q&A systems4 citations
TL;DR

This study evaluates large language models (LLMs) using multiple-choice questions (MCQs) based on a fictional medical gland, the Glianorex, to isolate pattern recognition from genuine clinical knowledge. Despite no prior exposure to the content, LLMs achieved ~67% accuracy across models, indicating MCQs primarily test statistical pattern matching rather than deep understanding, calling into question the validity of current benchmarks for assessing true medical reasoning in AI.

ABSTRACT

Large Language Models (LLMs) such as ChatGPT demonstrate significant potential in the medical domain and are often evaluated using multiple-choice questions (MCQs) modeled on exams like the USMLE. However, such benchmarks may overestimate true clinical understanding by rewarding pattern recognition and test-taking heuristics. To investigate this, we created a fictional medical benchmark centered on an imaginary organ, the Glianorex, allowing us to separate memorized knowledge from reasoning ability. We generated textbooks and MCQs in English and French using leading LLMs, then evaluated proprietary, open-source, and domain-specific models in a zero-shot setting. Despite the fictional content, models achieved an average score of 64%, while physicians scored only 27%. Fine-tuned medical models outperformed base models in English but not in French. Ablation and interpretability analyses revealed that models frequently relied on shallow cues, test-taking strategies, and hallucinated reasoning to identify the correct choice. These results suggest that standard MCQ-based evaluations may not effectively measure clinical reasoning and highlight the need for more robust, clinically meaningful assessment methods for LLMs.

Motivation & Objective

  • To investigate whether multiple-choice question (MCQ) benchmarks truly assess clinical knowledge and reasoning in large language models (LLMs) or merely test pattern recognition.
  • To evaluate the performance of diverse LLMs—open-source, proprietary, and fine-tuned—on a fictional medical domain with no real-world exposure.
  • To determine whether model size, architecture, or domain-specific fine-tuning improves performance on MCQs when assessing non-existent medical knowledge.
  • To challenge the validity of existing MCQ-based benchmarks in evaluating the true clinical understanding of LLMs in medicine.
  • To advocate for more robust evaluation methods beyond MCQs to better assess LLMs' clinical reasoning and knowledge.

Proposed method

  • Created a fictional medical gland, the Glianorex, and generated a comprehensive textbook on it using GPT-4 in both English and French.
  • Developed 264 multiple-choice questions per language based on the fictional textbook, ensuring no real medical content was involved.
  • Evaluated a diverse set of LLMs—including open-source, proprietary, and fine-tuned medical models—using a zero-shot setting with no fine-tuning on the fictional data.
  • Used high temperature sampling and multiple question generations per paragraph to reduce synthetic biases from the question-creation process.
  • Measured performance in both English and French to assess language-dependent pattern recognition effects.
  • Conducted statistical analysis to assess significance of performance differences across models and languages.

Experimental results

Research questions

  • RQ1To what extent do LLMs achieve high scores on MCQs based on entirely fictional medical knowledge, indicating pattern recognition over genuine understanding?
  • RQ2Do larger or fine-tuned LLMs perform significantly better on fictional MCQs compared to smaller or base models?
  • RQ3Is there a measurable performance difference between English and French versions of the same MCQs, suggesting language-specific pattern exploitation?
  • RQ4Do MCQ-based benchmarks effectively distinguish between superficial test-taking strategies and deep clinical reasoning in LLMs?
  • RQ5Can the high performance of LLMs on fictional MCQs undermine the validity of current benchmarks in assessing real medical knowledge?

Key findings

  • LLMs achieved an average accuracy of approximately 67% on multiple-choice questions based on a fictional medical gland, despite no prior exposure to the content.
  • Performance was nearly identical across models of varying sizes, architectures, and specializations, indicating a reliance on pattern recognition rather than deep knowledge.
  • Slight performance advantages were observed in English over French, suggesting language-specific statistical patterns may influence results.
  • Fine-tuned medical models showed marginal improvement over base models in English but no such gain in French, further highlighting the role of language and pattern matching.
  • The uniformly high performance across diverse models suggests that MCQs primarily assess test-taking strategies and statistical pattern matching, not clinical understanding.
  • The study concludes that current MCQ-based benchmarks may not be valid measures of true clinical knowledge or reasoning in LLMs, calling for more robust evaluation methods.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.