[논문 리뷰] A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
이 논문은 NeuroCognition을 제안하는데, 세 가지 신경심리학적 테스트(RPM, SWM, WCST)를 활용하는 다중 모달 벤치마크로서 표준 벤치마크를 넘어 LLM 인지 능력을 평가합니다. 텍스트에서의 강점과 이미지 및 복잡한 과제에서의 약점을 드러내고, 기존 벤치마크와의 상관관계를 보입니다.
Large language models (LLMs) exhibit a unified "general factor" of capability across 10 benchmarks, a finding confirmed by our factor analysis of 156 models, yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (maintenance and systematic search), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.
연구 동기 및 목표
- 확립된 신경심리학적 테스트를 LLM용으로 확장 가능한 다중 모달 벤치마크로 재목적화합니다.
- 현재 LLM이 추상적 추론, 작업 기억, 인지 유연성에서 어떻게 수행하는지 특징화합니다.
- 모달리티(텍스트 대 이미지) 및 과제 복잡도에 따라 모델 성능이 어떻게 달라지는지 평가합니다.
- 간단한 인간과 유사 전략(노트 필기, 단서 제시)이 LLM에 도움이 되는지 확인합니다.
- NeuroCognition과 표준 일반 역량 벤치마크 간의 관계를 탐구합니다.
제안 방법
- Raven’s Progressive Matrices(RPM)를 텍스트 및 이미지 형식의 추상 관계 추론에 맞게 적응합니다.
- 더 varying 난이도와 모달리티를 갖춘 유지 및 체계적 탐색을 측정하기 위해 Spatial Working Memory(SWM)를 적응합니다.
- 제어된 모호성 하에서 인지 유연성과 규칙 전환을 평가하기 위해 Wisconsin Card Sorting Test(WCST)를 적응합니다.
- 정확도, S_sw m, S_wcst 및 오류 유형 분석(합법적/ illegal, 박스 없음, 반복)을 포함한 성능 지표를 도입합니다.
- 패턴 힌트, 메모 등 인간과 흡사한 전략을 도입하여 인지적 오프로딩 효과를 평가합니다.
- 총 156개의 LLM과 10개의 벤치마크에 대한 요인 분석을 수행하여 일반 역량 인자(g)를 평가합니다.
실험 결과
연구 질문
- RQ1NeuroCognition로 측정된 일반 작업 수행을 넘어선 뚜렷한 인지 능력이 LLM에 존재합니까?
- RQ2모달리티(텍스트 대 이미지)와 과제 복잡도가 RPM, SWM, WCST에서 LLM의 성능에 어떤 영향을 줍니까?
- RQ3간단한 인간 유사 전략이 신경심리학적 과제에서 LLM의 성능을 개선합니까?
- RQ4NeuroCognition이 표준 일반-역량 벤치마크와 어떻게 관련이 있습니까?
- RQ5다양한 LLM 벤치마크에서 단일 차원 일반 요인(g)이 존재한다는 증거가 있습니까?
주요 결과
- LLMs는 텍스트 과제에서 강력한 성능을 보이나 이미지 과제와 과제 복잡도가 증가함에 따라 성능이 저하됩니다.
- 명시적 추론 강화가 보편적으로 이점이 되지 않으며, 경우에 따라 간단한 인간 유사 전략이 부분적 이득을 가져옵니다.
- NeuroCognition은 표준 벤치마크와 양의 상관관계가 있지만, 그것들로는 포착되지 않는 고유한 인지 능력을 포착합니다.
- 요인 분석은 10개의 벤치마크에서 약 75%의 분산을 설명하는 단일 잠재 일반 역량(g)을 보여주며, NeuroCognition은 서로 다른 인지 기본 요소를 목표로 합니다.
- 메모 필기 및 기타 인지-오프로딩 기법은 WCST에서 RAM 기반 과제보다 더 일관된 이득을 보이며 다양한 영향력을 보입니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.