Skip to main content
QUICK REVIEW

[論文レビュー] Emotional Intelligence of Large Language Models

Xuena Wang, Xueting Li|arXiv (Cornell University)|Jul 18, 2023
Topic Modeling被引用数 6
ひとこと要約

本研究では、現実的な社会的状況における感情理解(EU)に焦点を当て、大規模言語モデル(LLMs)の感情知能(EI)を評価するための新しい心理測定評価法を導入する。GPT-4はEQスコア117を達成し、人間参加者の89%を上回ったが、LLMsは人間とは異なった定性的な表象パターンを用いており、人間らしくないメカニズムによって人間水準のパフォーマンスを達成している可能性を示唆している。

ABSTRACT

Large Language Models (LLMs) have demonstrated remarkable abilities across numerous disciplines, primarily assessed through tasks in language generation, knowledge utilization, and complex reasoning. However, their alignment with human emotions and values, which is critical for real-world applications, has not been systematically evaluated. Here, we assessed LLMs' Emotional Intelligence (EI), encompassing emotion recognition, interpretation, and understanding, which is necessary for effective communication and social interactions. Specifically, we first developed a novel psychometric assessment focusing on Emotion Understanding (EU), a core component of EI, suitable for both humans and LLMs. This test requires evaluating complex emotions (e.g., surprised, joyful, puzzled, proud) in realistic scenarios (e.g., despite feeling underperformed, John surprisingly achieved a top score). With a reference frame constructed from over 500 adults, we tested a variety of mainstream LLMs. Most achieved above-average EQ scores, with GPT-4 exceeding 89% of human participants with an EQ of 117. Interestingly, a multivariate pattern analysis revealed that some LLMs apparently did not reply on the human-like mechanism to achieve human-level performance, as their representational patterns were qualitatively distinct from humans. In addition, we discussed the impact of factors such as model size, training method, and architecture on LLMs' EQ. In summary, our study presents one of the first psychometric evaluations of the human-like characteristics of LLMs, which may shed light on the future development of LLMs aiming for both high intellectual and emotional intelligence. Project website: https://emotional-intelligence.github.io/

研究の動機と目的

  • LLMsの感情知能(EI)を体系的に評価すること、特に社会的文脈における複雑な感情の理解能力に焦点を当てる。
  • 人間とLLMsの両方に適用可能な標準的で心理的妥当性のある感情理解(EU)テストを構築すること。
  • LLMsが人間水準のEIに到達する際、人間らしい認知的メカニズムを用いるのか、それとも代替的経路を用いるのかを調査すること。
  • モデルサイズ、トレーニング手法、アーキテクチャがLLMsの感情知能スコアに与える影響を分析すること。
  • 将来のLLMsにおける高い知的知能と感情知能の両立を促進するためのベンチマークを提供すること。

提案手法

  • 現実的な社会的状況を想定した複雑な感情(例:驚き、誇り、混乱)を含む、感情理解(EU)のための新しい心理的テストを開発した。
  • 500人を超える成人の回答を用いて人間の基準フレームを構築し、テストの補正と妥当性を検証した。
  • GPT-4を含む主流のLLMsにEUテストを実施し、人間のパフォーマンスと比較してEQスコアを測定した。
  • 多変量パターン分析を用いて、LLMsと人間の内部表象パターンを比較し、メカニズムの類似性を評価した。
  • モデルサイズ、トレーニング手法、アーキテクチャ設計がEIパフォーマンスに与える影響を定量化した。

実験結果

リサーチクエスチョン

  • RQ1標準化された心理的テストによって測定された場合、LLMsは現実的な社会的文脈における複雑な感情をどの程度理解できるか?
  • RQ2LLMsの感情知能スコアは、特にEQパーセンタイルの観点から人間と比較してどの程度か?
  • RQ3LLMsは人間が用いるメカニズムと類似した方法で高いEIに到達しているのか、それとも根本的に異なる表象パターンに依存しているのか?
  • RQ4モデルサイズ、トレーニング手法、アーキテクチャなどの要因が、LLMsの感情知能にどのように影響するか?

主な発見

  • GPT-4はEQスコア117を達成し、基準サンプルの89%の人間参加者を上回った。
  • テストされた大多数のLLMsは平均以上の中でのEQスコアを達成しており、人間のベンチマークと比較して感情理解のパフォーマンスが優れていることが示された。
  • 多変量パターン分析の結果、LLMsの感情理解のための内部表象は、人間とは定性的に異なることが判明し、人間らしくないメカニズムが関与している可能性を示唆した。
  • モデルサイズ、トレーニング手法、アーキテクチャはLLMsの感情知能に顕著な影響を与えることが判明したが、特定の効果はモデルごとに異なった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。