[论文解读] Evaluating Large Language Models in Theory of Mind Tasks
该研究在 40 个任务中,对 11 个大语言模型(LLMs)进行包含 640 条假信念提示的测试;GPT-4 解决了 75% 的任务,接近六岁儿童的表现,而较小的模型未解决任何任务,早期模型约达到 ~20%。
Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The battery included 640 prompts spread across 40 diverse tasks, each one including a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. To solve a single task, a model needed to correctly answer 16 prompts across all eight scenarios. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. We explore the potential interpretation of these findings, including the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.
研究动机与目标
- 评估大型语言模型在一组假信念任务中是否具备心智理论(ToM)能力。
- 比较来自小型/较旧模型到最先进模型的十一种 LLM 的表现。
- 量化当模型的语言能力提高时 ToM 类技能的出现程度。
提出的方法
- 使用自定义的假信念任务电池(跨 40 个任务的 640 条提示)。
- 每个任务包含一个假信念情景、三个匹配的真实信念对照,以及四种情景的相反版本。
- 一个模型要在八个情景中正确回答 16 条提示才能解决一个任务。
- 评估覆盖具有不同规模和架构的十一种 LLM。
- 结果与先前 ToM 研究中的人类基准进行比较说明。
- 代码和任务对复现开放(Colab)。
实验结果
研究问题
- RQ1十一种 LLM 在一组心智理论假信念任务中的表现如何?
- RQ2随着更新、更有能力的语言模型,ToM 表现是否得到改善?
- RQ3ToM 类能力是否作为语言建模进展的副产物而自发出现?
- RQ4模型表现与先前研究中的人类 ToM 基准相比如何?
主要发现
- 较小且较旧的 LLM 在该测试中没有解决任何任务。
- GPT-3-davinci-003(2022 年 11 月)和 ChatGPT-3.5-turbo(2023 年 3 月)解决了约 20% 的任务。
- ChatGPT-4(2023 年 6 月)解决了约 75% 的任务。
- GPT-4 的表现与过去六岁儿童研究中观察到的水平相匹配。
- 结果提出了一个可能性:随着 LLM 语言能力的提高,ToM 可能自发出现。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。