Skip to main content
QUICK REVIEW

[Paper Review] A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

Faiz Ghifari Haznitrama, Faeyza Rishad Ardi|arXiv (Cornell University)|Mar 3, 2026
Neurobiology of Language and Bilingualism0 citations
TL;DR

The paper proposes NeuroCognition, a multimodal benchmark using three neuropsychological tests (RPM, SWM, WCST) to evaluate LLM cognitive abilities beyond standard benchmarks, revealing strengths on text and weaknesses in image and complex tasks, with correlations to existing benchmarks.

ABSTRACT

Large language models (LLMs) exhibit a unified "general factor" of capability across 10 benchmarks, a finding confirmed by our factor analysis of 156 models, yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (maintenance and systematic search), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.

Motivation & Objective

  • Repurpose established neuropsychological tests into a scalable, multimodal benchmark for LLMs.
  • Characterize how current LLMs perform on abstract reasoning, working memory, and cognitive flexibility.
  • Assess how model performance varies with modality (text vs. image) and task complexity.
  • Examine whether simple human-like strategies (note-taking, hints) aid LLMs.
  • Explore the relationship between NeuroCognition and standard general-capability benchmarks.

Proposed method

  • Adapt Raven’s Progressive Matrices (RPM) for abstract relational reasoning in text and image formats.
  • Adapt Spatial Working Memory (SWM) to measure maintenance and systematic search with varying difficulty and modalities.
  • Adapt Wisconsin Card Sorting Test (WCST) to assess cognitive flexibility and rule-switching under controlled ambiguity.
  • Introduce performance metrics including accuracy, S_sw m, S_wcst, and error-type analyses (illegal, no-box, repeated).
  • Incorporate human-like strategies (pattern hints, note-taking) to evaluate cognitive offloading effects.
  • Perform factor analysis on 156 LLMs across 10 benchmarks to assess a general capability factor (g).

Experimental results

Research questions

  • RQ1Do LLMs exhibit distinct cognitive abilities beyond general task performance as measured by NeuroCognition?
  • RQ2How do modality (text vs. image) and task complexity affect LLM performance on RPM, SWM, and WCST?
  • RQ3Do simple human-like strategies improve LLM performance on neuropsychological tasks?
  • RQ4How does NeuroCognition relate to standard general-capability benchmarks?
  • RQ5Is there evidence of a unidimensional general factor (g) across diverse LLM benchmarks?

Key findings

  • LLMs perform strongly on text tasks but degrade on image tasks and as task complexity increases.
  • Explicit reasoning boosts are not universally beneficial; in some cases simple, human-like strategies yield partial gains.
  • NeuroCognition correlates positively with standard benchmarks yet captures distinct cognitive abilities beyond them.
  • Factor analysis reveals a single latent general capability (g) explaining ~75% of variance across 10 benchmarks, while NeuroCognition targets distinct cognitive primitives.
  • Note-taking and other cognitive-offloading techniques show variable impact, with more consistent gains in WCST than RAM-based tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.