[论文解读] SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
SOTOPIA 通过多样化、目标驱动的社会情境,引入了一个开放式的交互环境,用于评估语言智能体的社会智能。利用多维的 SOTOPIA-EVAL 框架,研究发现即使 GPT-4 在具有挑战性的社会任务中也表现不及人类,尤其在策略性沟通、社会常识和保密能力方面,凸显了当前大语言模型在社会推理能力上的显著差距。
Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
研究动机与目标
- 开发一个通用领域、交互式的环境,用于评估语言智能体的社会智能。
- 解决缺乏交互式、多维基准来评估人工智能智能体在复杂社会目标实现方面表现的问题。
- 识别并描述暴露当前大语言模型局限性的具有挑战性的社会情境。
- 使用全面的多维框架评估大语言模型和人类的表现。
- 评估基于大语言模型的判断与人类评估在社会智能指标上的一致性。
提出的方法
- SOTOPIA 通过组合随机化的情景、目标、角色、关系和智能体策略,生成多样化的人际互动片段。
- 智能体(包括基于大语言模型的和人类参与者)扮演角色,并通过言语、非言语和身体动作在多轮互动中进行交流。
- SOTOPIA-EVAL 从七个维度评估智能体表现:目标完成度、知识、信念、秘密、社会规范、财务和人际关系。
- 人类标注员对每个维度评分,同时使用 GPT-4 作为代理裁判以实现评估自动化。
- 在一般和困难情景子集上分析表现,并对模型与人类进行统计比较。
- 该框架利用人类评分的感知范围来验证 GPT-4 判断的一致性和可靠性。

实验结果
研究问题
- RQ1在开放式、交互式社会情境中,大语言模型与人类的表现如何比较?
- RQ2哪些社会情境在不同模型中持续具有挑战性,原因是什么?
- RQ3GPT-4 在多大程度上可作为人类判断社会智能的可靠代理?
- RQ4大语言模型在策略性沟通或保密能力等方面表现出哪些具体的社会推理失败?
- RQ5模型行为如何因对话伙伴的不同而变化,尤其是在高风险或复杂社会动态的情境中?
主要发现
- 在 SOTOPIA-hard 子集上,GPT-4 的目标完成率为 5.25(满分 10),显著低于人类表现(6.53,p < 0.05)。
- GPT-4 在社会规范(Soc: -0.38)和秘密(Sec: 0.00)维度得分较低,表明其倾向于违反社会规则或泄露机密信息。
- 尽管 GPT-4 的评估结果总体上处于人类估计的感知评分范围内,但在 Sec 和 Soc 维度上存在过度乐观的倾向。
- 尽管在知识和信念追踪方面表现良好(Kno: 7.63,Bel: 7.63),GPT-4 在策略性沟通和目标追求的持续性方面仍表现吃力。
- 在与他人互动时,人类参与者的人际关系维护能力(Rel: 0.93 vs. 0.65)和财务管理能力(Fin: 0.75 vs. 0.63)均优于 GPT-4。
- 在某些情况下,GPT-4 展现出创造性问题解决能力,但此类表现稀少且在不同情境中不一致。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。