Skip to main content
QUICK REVIEW

[論文レビュー] SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Xuhui Zhou, Hao Zhu|arXiv (Cornell University)|Oct 18, 2023
Topic ModelingComputer Science被引用数 3
ひとこと要約

SOTOPIA は、多様で目的志向の社会的シナリオを通じて、言語エージェントの社会的知性を評価する、オープンエンドでインタラクティブな環境を導入する。多面的な SOTOPIA-EVAL フレームワークを用いて、研究では GPT-4 ですら、戦略的コミュニケーション、社会的常識、秘密保持といった困難な社会的タスクにおいて人間を下回っていることが明らかになった。これは、現在の LLM が社会的推論能力に著しいギャップを抱えていることを示している。

ABSTRACT

Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.

研究の動機と目的

  • 言語エージェントの社会的知性を評価するための汎用ドメインでインタラクティブな環境を開発すること。
  • AI エージェントが複雑な社会的目的を達成する能力を評価するための、インタラクティブで多次元のベンチマークの不足に対処すること。
  • 現在の LLM に限界を露呈する、挑戦的な社会的シナリオを特定および特徴づけること。
  • 包括的で多次元的なフレームワークを用いて、LLM と人間のパフォーマンスを評価すること。
  • LLM による判断と人間の評価との間の整合性、特に社会的知性指標に関する整合性を評価すること。

提案手法

  • SOTOPIA は、ランダムに選択されたシナリオ、目的、登場人物、人間関係、エージェントのポリシーを組み合わせることで、多様な社会的相互作用エピソードを生成する。
  • LLM を用いたエージェントおよび人間参加者によるエージェントが、会話的・非言語的・身体的行動を用いて、複数ターンにわたって役割を演じる。
  • SOTOPIA-EVAL は、7 つの次元(目的達成、知識、信念、秘密、社会的規範、財務、関係)でエージェントのパフォーマンスを評価する。
  • 人間のアノテーターが各次元でエピソードをスコア付けし、GPT-4 をプロキシジャッジとして用いて自動評価を実施する。
  • 一般および困難なシナリオサブセットの両方でパフォーマンスを分析し、モデルと人間との間で統計的比較を実施する。
  • GPT-4 の判断の一貫性と信頼性を検証するために、人間が推定するスコア範囲をフレームワークが使用する。
Figure 2: Distribution of the difference between the scores given by humans and GPT-4.
Figure 2: Distribution of the difference between the scores given by humans and GPT-4.

実験結果

リサーチクエスチョン

  • RQ1LLM は、オープンエンドでインタラクティブな社会的シナリオにおいて、人間と比べてどのように性能を発揮するか?
  • RQ2どの社会的シナリオが、異なるモデルにおいて一貫して困難であるのか、その理由は何か?
  • RQ3GPT-4 は、社会的知性の評価において、人間の判断の信頼できるプロキシとしてどれほど有効に機能するか?
  • RQ4LLM が示す具体的な社会的推論の失敗、例えば戦略的コミュニケーションや秘密保持の分野での失敗はどのようなものか?
  • RQ5特に高リスクまたは複雑な社会的ダイナミクスにおいて、会話相手の違いに応じてモデルの行動はどのように変化するか?

主な発見

  • SOTOPIA-hard サブセットにおいて、GPT-4 は目的達成率が 10 中 5.25 にとどまり、人間のパフォーマンス(6.53、p < 0.05)を著しく下回った。
  • GPT-4 は社会的規範(Soc: -0.38)と秘密(Sec: 0.00)のスコアが低く、社会的ルールを破るか、機密情報を漏洩する傾向があることが示された。
  • GPT-4 の評価は一般的に人間が推定するスコア範囲内に収まったが、Sec および Soc の次元においては楽観的すぎた。
  • 知識および信念追跡(Kno: 7.63、Bel: 7.63)では高いパフォーマンスを示したが、戦略的コミュニケーションや目的達成への粘り強さには欠けていた。
  • 人間参加者は、他者との関係維持(Rel: 0.93 対 0.65)および財務管理(Fin: 0.75 対 0.63)において、GPT-4 を上回った。
  • 一部のケースでは GPT-4 が創造的な問題解決を示したが、そのようなケースは稀で、シナリオ間で一貫性がなかった。
Figure D.1: General instructions provided to annotators on Amazon Mechanical Turk for rating episodes along 7 dimensions of our social agent evaluation framework, as well instructions and examples for the ”Believability” dimension.
Figure D.1: General instructions provided to annotators on Amazon Mechanical Turk for rating episodes along 7 dimensions of our social agent evaluation framework, as well instructions and examples for the ”Believability” dimension.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。