[论文解读] Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs
本文研究了像 GPT-3 这类大型语言模型(LLMs)是否表现出神经理论心智,通过评估其在社会意图、情感和错误信念推理方面的能力。基于 SocialIQa 和 ToMi 基准测试,研究发现即使 GPT-3 的准确率也仅为 55% 和 60%,远低于人类表现,凸显了在数据、架构和训练范式限制下,尽管模型规模庞大,其社会智能仍存在根本性局限。
Social intelligence and Theory of Mind (ToM), i.e., the ability to reason about the different mental states, intents, and reactions of all people involved, allow humans to effectively navigate and understand everyday social interactions. As NLP systems are used in increasingly complex social situations, their ability to grasp social dynamics becomes crucial. In this work, we examine the open question of social intelligence and Theory of Mind in modern NLP systems from an empirical and theory-based perspective. We show that one of today's largest language models (GPT-3; Brown et al., 2020) lacks this kind of social intelligence out-of-the box, using two tasks: SocialIQa (Sap et al., 2019), which measures models' ability to understand intents and reactions of participants of social interactions, and ToMi (Le et al., 2019), which measures whether models can infer mental states and realities of participants of situations. Our results show that models struggle substantially at these Theory of Mind tasks, with well-below-human accuracies of 55% and 60% on SocialIQa and ToMi, respectively. To conclude, we draw on theories from pragmatics to contextualize this shortcoming of large language models, by examining the limitations stemming from their data, neural architecture, and training paradigms. Challenging the prevalent narrative that only scale is needed, we posit that person-centric NLP approaches might be more effective towards neural Theory of Mind. In our updated version, we also analyze newer instruction tuned and RLFH models for neural ToM. We find that even ChatGPT and GPT-4 do not display emergent Theory of Mind; strikingly even GPT-4 performs only 60% accuracy on the ToMi questions related to mental states and realities.
研究动机与目标
- 评估大型语言模型(LLMs)在现实社会推理任务中是否表现出社会智能和理论心智(ToM)能力。
- 评估最先进 LLMs(包括 GPT-3 及 GPT-3.5 和 GPT-4 等新型模型)在两项衡量社会常识和心智状态推理的基准任务上的表现。
- 通过分析训练数据、神经架构和训练范式,探究 LLM 在 ToM 任务中表现不佳的根本原因。
- 挑战当前普遍认为仅通过扩大模型规模即可实现神经理论心智的叙事,主张应采用以人为中心和交互式学习方法。
- 提供实证证据表明,即使经过微调和强化学习,当前 LLM 仍缺乏稳健的社会智能,并警示评估中可能出现的数据污染问题。
提出的方法
- 在 SocialIQa 基准测试上评估 GPT-3 和新型指令微调/RLHF 模型(如 GPT-3.5-Turbo、GPT-4),该测试衡量对社会意图、情感和反应的推理能力。
- 采用多种探测方法:零样本提示、少样本提示和多项选择探测,以评估不同提示策略下的模型表现。
- 在 ToMi 基准测试上评估模型,该测试受 Sally-Ann 错误信念测试启发,用于衡量对他人心智状态和现实的推理能力。
- 分析不同类型问题的表现差异,包括主要角色与次要参与者,以检测潜在的中心化或注意力偏差。
- 使用统计分析将模型表现与随机猜测(33%)和人类表现(ToMi 上 90–100%,SocialIQa 上 >85%)进行比较。
- 讨论数据泄露风险,指出 GPT-4 的训练数据可能包含 SocialIQa 和 ToMi 的测试样例,从而质疑高性能表现声明的有效性。
实验结果
研究问题
- RQ1大型语言模型(如 GPT-3)在日常互动中,能在多大程度上推理社会意图和情感反应?
- RQ2LLMs 能否准确推断他人的心智状态和错误信念,如 ToMi 基准测试所衡量的那样?
- RQ3与标准语言建模相比,指令微调或基于人类反馈的强化学习是否能提升 LLM 的社会推理能力?
- RQ4阻碍 LLM 实现人类水平理论心智的主要限制因素是什么——数据、架构还是训练范式?
- RQ5新型模型(如 GPT-3.5 和 GPT-4)观察到的性能提升,是源于真正的推理能力,还是基准数据集中的数据污染?
主要发现
- GPT-3 在衡量社会常识和情感智能的 SocialIQa 基准测试中准确率仅为 55%,落后于人类表现超过 30%。
- 在测试心智状态和错误信念推理的 ToMi 基准测试中,GPT-3 的准确率为 60%,仅比随机猜测(33%)高出 10%,而人类表现则为 90–100%。
- 指令微调和 RLHF 微调模型(如 GPT-3.5-RLHF)在 SocialIQa 上的表现提升至 55%,但仍远未达到人类水平的推理能力。
- GPT-4 在 SocialIQa 上得分为 79.3%,接近人类表现,但这一结果可能因训练集中存在潜在的数据泄露而被高估。
- 在 ToMi 测试中,GPT-3.5-Turbo 在心智状态问题上的准确率为 60%,略高于 GPT-4 的 59%,表明在该任务上规模增大并未带来一致性能提升。
- 所有模型在涉及次要角色的问题上表现显著劣于主要角色,表明其推理中仍存在持续的注意力或中心化偏差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。