[论文解读] Emotional Intelligence of Large Language Models
本研究提出了一种新颖的心理测量评估方法,用于衡量大语言模型(LLMs)的情绪智力(EI),重点关注在真实社交情境中的情绪理解(EU)。GPT-4 的情商(EQ)得分为 117,超过 89% 的人类参与者,尽管 LLMs 使用了与人类不同的表征模式,表明其达到人类水平表现的机制并非类人。
Large Language Models (LLMs) have demonstrated remarkable abilities across numerous disciplines, primarily assessed through tasks in language generation, knowledge utilization, and complex reasoning. However, their alignment with human emotions and values, which is critical for real-world applications, has not been systematically evaluated. Here, we assessed LLMs' Emotional Intelligence (EI), encompassing emotion recognition, interpretation, and understanding, which is necessary for effective communication and social interactions. Specifically, we first developed a novel psychometric assessment focusing on Emotion Understanding (EU), a core component of EI, suitable for both humans and LLMs. This test requires evaluating complex emotions (e.g., surprised, joyful, puzzled, proud) in realistic scenarios (e.g., despite feeling underperformed, John surprisingly achieved a top score). With a reference frame constructed from over 500 adults, we tested a variety of mainstream LLMs. Most achieved above-average EQ scores, with GPT-4 exceeding 89% of human participants with an EQ of 117. Interestingly, a multivariate pattern analysis revealed that some LLMs apparently did not reply on the human-like mechanism to achieve human-level performance, as their representational patterns were qualitatively distinct from humans. In addition, we discussed the impact of factors such as model size, training method, and architecture on LLMs' EQ. In summary, our study presents one of the first psychometric evaluations of the human-like characteristics of LLMs, which may shed light on the future development of LLMs aiming for both high intellectual and emotional intelligence. Project website: https://emotional-intelligence.github.io/
研究动机与目标
- 系统评估 LLM 的情绪智力(EI),特别是其在社交情境中理解复杂情绪的能力。
- 开发一种标准化、具有心理测量学有效性的情绪理解(EU)测试,适用于人类和 LLM。
- 探究 LLM 是否通过类人认知机制或替代路径实现人类水平的 EI。
- 分析模型规模、训练方法和架构对 LLM 情绪智力得分的影响。
- 为未来开发兼具高智力与高情绪智力的 LLM 提供基准。
提出的方法
- 基于包含复杂情绪(如惊讶、自豪、困惑)的真实社交情境,开发了一种新颖的心理测量 EU 测试。
- 通过超过 500 名成人的回答构建人类参考框架,以校准和验证测试。
- 对一系列主流 LLM(包括 GPT-4)实施 EU 测试,以衡量其相对于人类表现的 EQ 得分。
- 应用多变量模式分析,比较 LLM 与人类的内部表征模式,以评估机制相似性。
- 量化模型规模、训练方法和架构设计对 EI 表现的影响。
实验结果
研究问题
- RQ1在标准化心理测量测试下,LLMs 在真实社交情境中理解复杂情绪的程度如何?
- RQ2LLMs 的情绪智力得分与人类相比如何,特别是在 EQ 百分位数上的差异?
- RQ3LLMs 是否通过与人类相似的机制实现高 EI,还是依赖于根本不同的表征模式?
- RQ4模型规模、训练方法和架构等因素如何影响 LLM 的情绪智力?
主要发现
- GPT-4 的 EQ 得分为 117,使其在参考样本中超过 89% 的人类参与者。
- 大多数测试的 LLM 在 EQ 得分上均高于平均水平,表明其在情绪理解方面相对于人类基准表现强劲。
- 多变量模式分析显示,LLMs 在情绪理解方面的内部表征与人类存在定性差异,表明其机制并非类人。
- 模型规模、训练方法和架构被发现显著影响 LLM 的情绪智力,尽管具体影响在不同模型间存在差异。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。