[Paper Review] Evaluating the Deductive Competence of Large Language Models
This study evaluates the deductive reasoning abilities of large language models (LLMs) using the Wason selection task, testing how content type (social vs. non-social rules) and presentation format affect performance. Despite training on vast text corpora, LLMs show low accuracy (10-20%) on abstract problems and only modest gains on social-rule problems, with performance influenced by unexpected, non-human-like interactions between content and format, indicating emergent reasoning biases distinct from human cognition.
The development of highly fluent large language models (LLMs) has prompted increased interest in assessing their reasoning and problem-solving capabilities. We investigate whether several LLMs can solve a classic type of deductive reasoning problem from the cognitive science literature. The tested LLMs have limited abilities to solve these problems in their conventional form. We performed follow up experiments to investigate if changes to the presentation format and content improve model performance. We do find performance differences between conditions; however, they do not improve overall performance. Moreover, we find that performance interacts with presentation format and content in unexpected ways that differ from human performance. Overall, our results suggest that LLMs have unique reasoning biases that are only partially predicted from human reasoning performance and the human-generated language corpora that informs them.
Motivation & Objective
- To assess whether LLMs can solve classic deductive reasoning problems from cognitive science, particularly the Wason selection task.
- To investigate whether social-rule content improves LLM performance, as it does in humans.
- To examine the impact of different presentation formats on LLM reasoning performance.
- To determine whether performance differences are consistent across diverse LLM architectures and training data.
- To identify emergent reasoning biases in LLMs that deviate from human cognitive patterns.
Proposed method
- Administered the Wason selection task to multiple LLMs using standardized, controlled problem sets with varying content (social, realistic, arbitrary) and presentation formats (classic, shuffled, rephrased).
- Designed experiments with fixed problem content and varied formatting to isolate the effects of presentation on model performance.
- Measured model responses for correctness, focusing on selection of logically valid cards (modus ponens and modus tollens).
- Compared performance across LLMs with different architectures, training data, and objectives to assess consistency of findings.
- Analyzed response patterns for antecedent vs. consequent card selection to detect reasoning biases.
- Used controlled, zero-shot prompting to avoid fine-tuning, ensuring evaluation of general reasoning ability.

Experimental results
Research questions
- RQ1Does the inclusion of social-rule content improve deductive reasoning performance in LLMs, as it does in humans?
- RQ2How do different presentation formats (e.g., classic vs. shuffled) affect LLM performance on Wason tasks?
- RQ3Are there consistent, non-human-like interactions between content type and presentation format in LLM reasoning?
- RQ4To what extent do LLMs replicate human reasoning patterns in conditional inference tasks?
- RQ5Do reasoning biases in LLMs remain consistent across different models despite architectural and training differences?
Key findings
- LLMs exhibit low performance on abstract Wason problems, with accuracy around 10-20%, consistent with human baseline performance.
- Performance on social-rule problems improves slightly but does not reach human-level accuracy (70%+), indicating limited benefit from content familiarity.
- Presentation format significantly affects performance, with the classic format yielding the highest accuracy, though improvements are marginal.
- Unexpected, non-human-like interactions between content and format were observed, suggesting reasoning biases not predicted by human cognitive models.
- LLMs show inconsistent card selection behavior, particularly underperforming on antecedent card selection, especially in non-social and realistic rule conditions.
- Despite differences in training data and architecture, performance patterns and reasoning biases are remarkably consistent across LLMs, indicating emergent, shared cognitive biases.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.