[论文解读] Language Writ Large: LLMs, ChatGPT, Grounding, Meaning and Understanding
本文认为像 ChatGPT 这样的 LLMs 并不真正理解;它讨论在 LLM 规模上的良性偏见、语义定位(grounding)缺口,并提出关于间接语言语义定位、循环性以及其他因素如何塑造其能力的直觉假设,以与 ChatGPT-4 的对话为框架。
Apart from what (little) OpenAI may be concealing from us, we all know (roughly) how ChatGPT works (its huge text database, its statistics, its vector representations, and their huge number of parameters, its next-word training, and so on). But none of us can say (hand on heart) that we are not surprised by what ChatGPT has proved to be able to do with these resources. This has even driven some of us to conclude that ChatGPT actually understands. It is not true that it understands. But it is also not true that we understand how it can do what it can do. I will suggest some hunches about benign biases: convergent constraints that emerge at LLM scale that may be helping ChatGPT do so much better than we would have expected. These biases are inherent in the nature of language itself, at LLM scale, and they are closely linked to what it is that ChatGPT lacks, which is direct sensorimotor grounding to connect its words to their referents and its propositions to their meanings. These convergent biases are related to (1) the parasitism of indirect verbal grounding on direct sensorimotor grounding, (2) the circularity of verbal definition, (3) the mirroring of language production and comprehension, (4) iconicity in propositions at LLM scale, (5) computational counterparts of human categorical perception in category learning by neural nets, and perhaps also (6) a conjecture by Chomsky about the laws of thought. The exposition will be in the form of a dialogue with ChatGPT-4.
研究动机与目标
- 评估像 ChatGPT 这样的 LLM 是否具备真正的理解或意义。
- 识别在 LLM 规模下出现、影响性能的趋同偏见。
- 考察缺乏直接的感知运动 grounding 如何影响指称对象和命题。
- 提出关于大型语言模型中的 grounding、定义和表征的假设。
提出的方法
- 以 ChatGPT-4 的对话形式呈现以探讨该主题。
- 讨论可能在 LLM 规模下出现的一组趋同偏见。
- 概述间接语言 grounding、定义的循环性以及知觉分类之间的联系。
- 将论点与更广泛的 grounding 和语言理解理论联系起来。
实验结果
研究问题
- RQ1像 ChatGPT 一样的 LLM 真正理解其输出的含义吗?
- RQ2在 LLM 规模下会出现哪些趋同偏见,从而提升或限制性能?
- RQ3缺乏直接的感知运动 grounding 如何影响 LLMs 中的指称对象、意义和命题?
- RQ4哪些理论视角(如 Chomsky、类别化、象性等理论视角)能揭示 LLM 的能力与局限?
主要发现
- LLMs 可以完成令人印象深刻的任务,但并不具备真正的理解或感知运动经验中的 grounding。
- 规模上的趋同偏见可能源于语言数据的结构与分布以及模型架构。
- 间接的语言 grounding 似乎在某种程度上替代了直接的感知运动 grounding,影响指称连接。
- 语言定义的循环性以及产出与理解之间的镜像关系与 LLM 行为相关。
- 象性和类别学习原则可能对应于 LLM 规模的表征,为性能提供部分解释。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。