[论文解读] Dissociating language and thought in large language models
本文区分形式性语言能力与功能性语言能力,表明大型语言模型在语言的形式方面表现出色,但在没有专门微调或外部模块的情况下,面向世界的功能性推理落后。
Large Language Models (LLMs) have come closest among all models to date to mastering human language, yet opinions about their linguistic and cognitive capabilities remain split. Here, we evaluate LLMs using a distinction between formal linguistic competence -- knowledge of linguistic rules and patterns -- and functional linguistic competence -- understanding and using language in the world. We ground this distinction in human neuroscience, which has shown that formal and functional competence rely on different neural mechanisms. Although LLMs are surprisingly good at formal competence, their performance on functional competence tasks remains spotty and often requires specialized fine-tuning and/or coupling with external modules. We posit that models that use language in human-like ways would need to master both of these competence types, which, in turn, could require the emergence of mechanisms specialized for formal linguistic competence, distinct from functional competence.
研究动机与目标
- 推动并形式化区分形式性语言能力(规则与模式)与功能性语言能力(在世界中使用语言)。
- 评估当代 LLM 在大规模下是否实现形式性语言能力,并识别在推理、世界知识、情境建模和社会认知等领域的功能性能力差距。
- 将形式性/功能性区分置于人类神经科学基础,以解读 LLM 能力与局限。
- 讨论对构建和评估未来语言模型和通用人工智能(AGI)的影响。
提出的方法
- 综述认知科学和神经科学中关于语言与思维分离的现有证据。
- 评估 LLM 在形式性语言任务上的表现(例如层级结构、长距离依赖),使用如 BLiMP 与 SyntaxGym 的基准测试。
- 分析规模化与微调(如 RLHF)在各领域对功能性能力的影响。
- 提供机制性与探针视角,以解读内部表示是否编码抽象语言结构。
- 将模型行为与人类神经结构的比较基础,以将语言处理与非语言认知区分开来。

实验结果
研究问题
- RQ1LLMs 是否展示出可与人类规则和层级结构理解相当的形式性语言能力?
- RQ2LLMs 在多大程度上表现出如现实世界推理、世界知识和社会认知等功能性语言能力?
- RQ3规模化、微调以及外部模块的增强相对于形式性能力对功能性能力有何影响?
- RQ4来自神经科学和认知科学的哪些证据支持将语言处理与一般思维在人类 LLMs 中区分?
主要发现
- LLMs 展示出强大的形式性语言能力,随着数据与规模的增加,掌握了许多复杂的语言现象。
- 如 BLiMP 和 SyntaxGym 的基准测试显示,在语法性与句法依赖方面达到高水平但尚未达到人类水平的表现。
- LLMs 学习抽象和层级结构,包括长距离主谓一致和非局部依赖。
- 功能性能力仍然零散且高度依赖任务/领域,通常需要微调或与外部模块耦合以实现世界知识、推理或社会认知任务。
- 形式性能力随训练数据显著提升,而功能性能力的提升不那么显著,且依赖超越仅靠规模的数据的专门方法。
- 人类语言网络可与非语言认知区分,这支持这样的观点:尽管具备语言能力,语言模型可能无法完整捕捉人类思维。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。