[論文レビュー] Do LLM Agents Exhibit Social Behavior?
本論文は、行動経済学のゲームを用いて、大規模言語モデル(LLMs)における社会的行動を体系的に評価するための確率的フレームワークであるState-Understanding-Value-Action(SUVA)を導入する。SUVAは発話に基づく推論を分析し、意思決定を予測する。その結果、大多数のLLMは公平性や報酬の対応といった対人的行動を示しており、特に能力の高いモデルでは団体への帰属意識の影響が顕著になる。また、利他行動に関する推論の参照は対人的行動を増加させ、自己利益の参照はそれを減少させる。
As LLMs increasingly take on roles in human-AI interactions and autonomous AI systems, understanding their social behavior becomes important for informed use and continuous improvement. However, their behaviors in social interactions with humans and other agents, as well as the mechanisms shaping their responses, remain underexplored. To address this gap, we introduce a novel probabilistic framework, State-Understanding-Value-Action (SUVA), to systematically analyze LLM responses in social contexts based on their textual outputs (i.e., utterances). Using canonical behavioral economics games and social preference concepts relatable to LLM users, SUVA assesses LLMs' social behavior through both their final decisions and the response generation processes leading to those decisions. Our analysis of eight LLMs -- including two GPT, four LLaMA, and two Mistral models -- suggests that most models do not generate decisions aligned solely with self-interest; instead, they often produce responses that reflect social welfare considerations and display patterns consistent with direct and indirect reciprocity. Additionally, higher-capacity models more frequently display group identity effects. The SUVA framework also provides explainable tools -- including tree-based visualizations and probabilistic dependency analysis -- to elucidate how factors in LLMs' utterance-based reasoning influence their decisions. We demonstrate that utterance-based reasoning reliably predicts LLMs' final actions; references to altruism, fairness, and cooperation in the reasoning increase the likelihood of prosocial actions, while mentions of self-interest and competition reduce them. Overall, our framework enables practitioners to assess LLMs for applications involving social interactions, and provides researchers with a structured method to interpret how LLM behavior arises from utterance-based reasoning.
研究の動機と目的
- 人間-AI相互作用およびマルチエージェント相互作用におけるLLMの社会的行動を体系的に評価するフレームワークの不足を解消すること。
- 特に公平性、報酬の対応、自己利益に関して、LLMが社会的文脈でどのように意思決定を行うかを理解すること。
- 発話に基づく推論が最終的意思決定に与える影響を、透明かつ解釈可能な方法で評価する手法を開発すること。
- 実世界のAI導入において、組織の価値観に適合したモデル選定を支援すること。
- 予測可能で現実的な社会的行動を保証することで、LLMをエージェントベースモデリングに信頼性を持って統合できること。
提案手法
- BDI心理学を模倣した確率的モデルであるSUVAフレームワークを提唱し、社会的意思決定におけるLLMの反応を分析する。
- 代表的な行動経済学のゲーム(例:独裁者ゲーム)を用いて、LLMの意思決定と推論を誘発・評価する。
- 木構造ベースの可視化と確率的依存関係分析を用い、発話内容が最終的意思決定にどのように影響するかをマッピングする。
- Chain-of-Thought(CoT)推論における利他行動、公平性、競争といった社会的好みを定量化し、結果を予測する。
- 統計的モデリングを用いて、推論内容に基づくLLM意思決定の予測可能性を評価し、次トークン予測を確率的意思決定経路として扱う。
- CoTの予測可能性分析を通じたフレームワークの妥当性を検証し、推論内容が最終行動を信頼性を持って予測できることを示した。
実験結果
リサーチクエスチョン
- RQ1LLMエージェントは、社会的意思決定の文脈において、公平性や利他行動といった対人的行動を示すか?
- RQ2発話に基づく推論は、LLMの社会的相互作用における最終的意思決定にどのように影響するか?
- RQ3モデルアーキテクチャーや能力が、LLMの反応における社会的好みの表現に及ぼす影響の程度はどの程度か?
- RQ4公平性や自己利益に関する参照といった推論内容が、LLMの最終的意思決定を信頼性を持って予測できるか?
- RQ5GPT、LLaMA、Mistralなどの異なるLLMファミリーは、社会的行動および推論パターンにおいてどのように異なるか?
主な発見
- 大多数のLLMは自己利益に基づいて行動するのではなく、社会的福祉や対人的価値を反映した反応を頻繁に生成する。
- 能力の高いLLMはより一貫して団体への帰属意識の影響を示しており、能力が社会的行動の表現に影響を与えることが示唆される。
- GPTおよびMistralモデルでは、モデル容量の増加に伴い自己利益の表現が減少するが、LLaMAモデルでは逆の傾向が見られる。
- 発話に基づく推論は最終的意思決定を信頼性高く予測できる:利他行動、公平性、協力に関する参照は、対人的行動の可能性を高める。
- 推論における自己利益や競争の言及は、対人的行動の確率を低下させ、推論内容と行動との間に因果関係があることを示している。
- SUVAフレームワークは、木構造ベースの構造と確率的依存関係分析を通じて意思決定経路を効果的に可視化・説明でき、LLM行動の解釈可能性を向上させた。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。