[论文解读] A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
论文提出 NeuroCognition,一种多模态基准,使用三种神经心理测试(RPM、SWM、WCST)来评估语言模型在超越标准基准的认知能力,揭示文本任务上的优势以及图像与复杂任务上的弱点,并与现有基准相关性分析。
Large language models (LLMs) exhibit a unified "general factor" of capability across 10 benchmarks, a finding confirmed by our factor analysis of 156 models, yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (maintenance and systematic search), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.
研究动机与目标
- 将已确立的神经心理测试重新用于可扩展的、面向多模态的大语言模型基准。
- 表征当前大语言模型在抽象推理、工作记忆和认知灵活性方面的表现。
- 评估模型在模态(文本 vs 图像)与任务复杂度上的表现差异。
- 检验简单的人类式策略(记笔记、提示)是否有助于大语言模型。
- 探索 NeuroCognition 与标准通用能力基准之间的关系。
提出的方法
- 将 Raven’s Progressive Matrices (RPM) 改编为文本和图像格式的抽象关系推理任务。
- 将 Spatial Working Memory (SWM) 改编为在不同难度和模态下的维护与有序搜索测量。
- 将 Wisconsin Card Sorting Test (WCST) 改编为在受控模糊环境中评估认知灵活性与规则切换。
- 引入包括准确度、S_sw m、S_wcst 与错误类型分析(非法、无盒、重复)的性能度量。
- 引入类人类策略(模式提示、记笔记)以评估认知卸载效应。
- 对156个 LLM 在10个基准上的因子分析,以评估通用能力因子(g)。
实验结果
研究问题
- RQ1LLMs 是否表现出超越 NeuroCognition 测量的一般任务表现的不同认知能力?
- RQ2模态(文本 vs 图像)和任务复杂度如何影响 LLM 在 RPM、SWM、WCST 上的表现?
- RQ3简单的人类式策略是否能改善 LLM 在神经心理测试中的表现?
- RQ4NeuroCognition 与标准通用能力基准之间的关系如何?
- RQ5是否存在跨越多样化基准的一维通用因子(g)的证据?
主要发现
- LLMs 在文本任务上表现强劲,但在图像任务和任务复杂度提升时表现下降。
- 明确的推理增强并不普遍有利;在某些情况下,简单的人类式策略可带来部分收益。
- NeuroCognition 与标准基准呈正相关,但也捕捉到超出它们的独特认知能力。
- 因子分析揭示一个单一潜在的通用能力(g)能够解释10个基准大约75%的方差,而 NeuroCognition 针对的是不同的认知基元。
- 记笔记等认知卸载技巧的影响各异,在 WCST 上比在基于 RAM 的任务上具有更稳定的收益。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。