Skip to main content
QUICK REVIEW

[论文解读] Evaluating Large Language Models in Theory of Mind Tasks

Michał Kosiński|arXiv (Cornell University)|Feb 4, 2023
Topic Modeling被引用 132
一句话总结

该研究在 40 个任务中,对 11 个大语言模型(LLMs)进行包含 640 条假信念提示的测试;GPT-4 解决了 75% 的任务,接近六岁儿童的表现,而较小的模型未解决任何任务,早期模型约达到 ~20%。

ABSTRACT

Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The battery included 640 prompts spread across 40 diverse tasks, each one including a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. To solve a single task, a model needed to correctly answer 16 prompts across all eight scenarios. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. We explore the potential interpretation of these findings, including the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.

研究动机与目标

  • 评估大型语言模型在一组假信念任务中是否具备心智理论(ToM)能力。
  • 比较来自小型/较旧模型到最先进模型的十一种 LLM 的表现。
  • 量化当模型的语言能力提高时 ToM 类技能的出现程度。

提出的方法

  • 使用自定义的假信念任务电池(跨 40 个任务的 640 条提示)。
  • 每个任务包含一个假信念情景、三个匹配的真实信念对照,以及四种情景的相反版本。
  • 一个模型要在八个情景中正确回答 16 条提示才能解决一个任务。
  • 评估覆盖具有不同规模和架构的十一种 LLM。
  • 结果与先前 ToM 研究中的人类基准进行比较说明。
  • 代码和任务对复现开放(Colab)。

实验结果

研究问题

  • RQ1十一种 LLM 在一组心智理论假信念任务中的表现如何?
  • RQ2随着更新、更有能力的语言模型,ToM 表现是否得到改善?
  • RQ3ToM 类能力是否作为语言建模进展的副产物而自发出现?
  • RQ4模型表现与先前研究中的人类 ToM 基准相比如何?

主要发现

  • 较小且较旧的 LLM 在该测试中没有解决任何任务。
  • GPT-3-davinci-003(2022 年 11 月)和 ChatGPT-3.5-turbo(2023 年 3 月)解决了约 20% 的任务。
  • ChatGPT-4(2023 年 6 月)解决了约 75% 的任务。
  • GPT-4 的表现与过去六岁儿童研究中观察到的水平相匹配。
  • 结果提出了一个可能性:随着 LLM 语言能力的提高,ToM 可能自发出现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。