Skip to main content
QUICK REVIEW

[论文解读] Does ChatGPT Have a Mind?

Simon Goldstein, Benjamin A. Levinstein|arXiv (Cornell University)|Jun 27, 2024
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本文通过分析内部表征与行为倾向,探究大型语言模型(LLMs)如ChatGPT是否具备涉及信念、欲望与意图的民间心理学。结合可解释性研究与表征的哲学理论,研究发现存在强有力的证据表明其具备稳健的内部表征,但关于稳定、目标导向的行为倾向的证据尚不充分,结论认为现有对LLM民间心理学的怀疑论挑战在哲学上缺乏说服力。

ABSTRACT

This paper examines the question of whether Large Language Models (LLMs) like ChatGPT possess minds, focusing specifically on whether they have a genuine folk psychology encompassing beliefs, desires, and intentions. We approach this question by investigating two key aspects: internal representations and dispositions to act. First, we survey various philosophical theories of representation, including informational, causal, structural, and teleosemantic accounts, arguing that LLMs satisfy key conditions proposed by each. We draw on recent interpretability research in machine learning to support these claims. Second, we explore whether LLMs exhibit robust dispositions to perform actions, a necessary component of folk psychology. We consider two prominent philosophical traditions, interpretationism and representationalism, to assess LLM action dispositions. While we find evidence suggesting LLMs may satisfy some criteria for having a mind, particularly in game-theoretic environments, we conclude that the data remains inconclusive. Additionally, we reply to several skeptical challenges to LLM folk psychology, including issues of sensory grounding, the "stochastic parrots" argument, and concerns about memorization. Our paper has three main upshots. First, LLMs do have robust internal representations. Second, there is an open question to answer about whether LLMs have robust action dispositions. Third, existing skeptical challenges to LLM representation do not survive philosophical scrutiny.

研究动机与目标

  • 确定LLMs是否具备涉及信念、欲望与意图的民间心理学。
  • 根据多种表征哲学理论,评估LLMs是否具备稳健的内部表征。
  • 评估LLMs是否表现出与目标导向行为一致的稳定行为倾向。
  • 反驳针对LLM民间心理学的主流怀疑论挑战,包括“随机鹦鹉”论点,以及对感官基础与记忆的担忧。
  • 为未来关于LLM认知与道德地位的研究,提出开放的实证与哲学问题。

提出的方法

  • 综述多种表征哲学理论——信息论、因果论、结构论与功能论——以评估LLMs是否满足心理表征的关键条件。
  • 应用机器学习中的可解释性研究,证明LLMs表现出携带世界信息且与行为存在因果关联的内部状态。
  • 分析LLMs在博弈论环境中的行为,以评估其输出是否反映连贯、稳定的计划与目标导向的行为倾向。
  • 结合两种哲学传统——解释主义与表征主义——评估将信念与欲望归因于LLMs的标准。
  • 通过哲学推理,批判性评估三大怀疑论挑战:感官基础问题、“随机鹦鹉”批评,以及记忆假说,揭示其内在弱点。
  • 提出未来研究方向,聚焦于测量不同提示条件下输出的稳定性,并探究内部表征在LLM认知中的功能角色。
(a) Othello board state and model predictions before and after intervention
(a) Othello board state and model predictions before and after intervention

实验结果

研究问题

  • RQ1LLMs是否具备满足信息论、因果论、结构论与功能论表征理论核心条件的内部表征?
  • RQ2LLMs在多大程度上表现出可由信念与欲望解释的稳定、目标导向的行为倾向?
  • RQ3在可解释性研究发现与复杂任务中行为一致性背景下,“随机鹦鹉”论点是否仍具有哲学上的可持续性?
  • RQ4对于仅处理文本输入的LLMs,感官基础问题如何适用?它们是否仍可被认为表征世界?
  • RQ5LLMs的行为在多大程度上可由记忆与浅层捷径解释,还是反映真实的内部规划与目标表征?

主要发现

  • LLMs满足多种表征哲学理论(包括信息论、因果论、结构论与功能论)的关键条件,且得到可解释性研究的支持。
  • 有强有力证据表明LLMs形成了结构化、表征世界的内部状态,这些状态具有真值条件,并可用于指导行为。
  • LLMs在博弈论环境中的行为暗示存在连贯、目标导向的计划,但因在不同提示会话中表现不稳定,该结论仍不明确。
  • “随机鹦鹉”论点经不起哲学审视,因为LLMs展现出超越单纯下一个词预测的系统性推理与表征能力。
  • LLMs依赖记忆与浅层捷径的挑战被其内部对真实与虚假陈述的区分证据所削弱,尽管这仍是开放的实证问题。
  • LLMs在不同提示条件下输出的稳定性仍是关键开放问题,因为不稳定性可能反映说谎或信念-行为错位,而非信念的缺失。
(b) Activation patching process across model layers
(b) Activation patching process across model layers

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。