[Paper Review] Dissociating language and thought in large language models
The paper distinguishes formal versus functional linguistic competence, showing LLMs excel at formal aspects of language but lag in functional, world-directed reasoning without specialized fine-tuning or external modules.
Large Language Models (LLMs) have come closest among all models to date to mastering human language, yet opinions about their linguistic and cognitive capabilities remain split. Here, we evaluate LLMs using a distinction between formal linguistic competence -- knowledge of linguistic rules and patterns -- and functional linguistic competence -- understanding and using language in the world. We ground this distinction in human neuroscience, which has shown that formal and functional competence rely on different neural mechanisms. Although LLMs are surprisingly good at formal competence, their performance on functional competence tasks remains spotty and often requires specialized fine-tuning and/or coupling with external modules. We posit that models that use language in human-like ways would need to master both of these competence types, which, in turn, could require the emergence of mechanisms specialized for formal linguistic competence, distinct from functional competence.
Motivation & Objective
- Motivate and formalize a distinction between formal linguistic competence (rules and patterns) and functional linguistic competence (using language in the world).
- Assess whether contemporary LLMs achieve formal linguistic competence at scale and identify gaps in functional competence across domains like reasoning, world knowledge, situation modeling, and social cognition.
- Ground the formal/functional distinction in human neuroscience to interpret LLM capabilities and limitations.
- Discuss implications for building and evaluating future language models and AGI.
Proposed method
- Review existing evidence from cognitive science and neuroscience on language vs. thought dissociation.
- Evaluate LLMs’ performance on formal linguistic tasks (e.g., hierarchical structure, long-distance dependencies) using benchmarks like BLiMP and SyntaxGym."
- Analyze how scaling and fine-tuning (e.g., RLHF) affect functional competence across domains.
- Provide mechanistic and probing perspectives to interpret whether internal representations encode abstract linguistic structure.
- Ground comparisons between model behavior and human neural architecture to separate language processing from non-linguistic cognition.

Experimental results
Research questions
- RQ1Do LLMs demonstrate formal linguistic competence comparable to human-like rules and hierarchical structure understanding?
- RQ2To what extent do LLMs exhibit functional linguistic competence such as real-world reasoning, world knowledge, and social cognition?
- RQ3How do scaling, fine-tuning, and augmentation with external modules impact functional competence relative to formal competence?
- RQ4What evidence from neuroscience and cognitive science supports dissociating language processing from general thought in LLMs?
Key findings
- LLMs show strong formal linguistic competence, mastering many complex linguistic phenomena as data and scale increase.
- Benchmarks like BLiMP and SyntaxGym reveal high but not perfect human-level performance on grammaticality and syntactic dependencies.
- LLMs learn abstractions and hierarchical structures, including long-distance subject–verb agreement and non-local dependencies.
- Functional competence remains patchy and highly task/domain dependent, often requiring fine-tuning or coupling with external modules to perform world knowledge, reasoning, or social-cognitive tasks.
- Formal competence improves substantially with training data, whereas functional competence gains are less dramatic and rely on specialized methods beyond scaling alone.
- The language network in humans is dissociable from non-linguistic cognition, supporting the view that language models may not fully capture human thought despite linguistic prowess.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.