[论文解读] Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
该论文表明,大型语言模型(LLMs)可以生成合成数据,用于训练模型以预测语言型说服博弈中的人类选择,当可用的 LLM 生成数据足够多时,其预测准确率高于使用真实人类数据训练的模型。该方法使用具有不同人格的 LLM 代理来模拟策略性互动,其表现优于基于人类数据的基线模型,甚至在关键设置中超越了人类数据模型。
Human choice prediction in economic contexts is crucial for applications in marketing, finance, public policy, and more. This task, however, is often constrained by the difficulties in acquiring human choice data. With most experimental economics studies focusing on simple choice settings, the AI community has explored whether LLMs can substitute for humans in these predictions and examined more complex experimental economics settings. However, a key question remains: can LLMs generate training data for human choice prediction? We explore this in language-based persuasion games, a complex economic setting involving natural language in strategic interactions. Our experiments show that models trained on LLM-generated data can effectively predict human behavior in these games and even outperform models trained on actual human data. Beyond data generation, we investigate the dual role of LLMs as both data generators and predictors, introducing a comprehensive empirical study on the effectiveness of utilizing LLMs for data generation, human choice prediction, or both. We then utilize our choice prediction framework to analyze how strategic factors shape decision-making, showing that interaction history (rather than linguistic sentiment alone) plays a key role in predicting human decision-making in repeated interactions. Particularly, when LLMs capture history-dependent decision patterns similarly to humans, their predictive success improves substantially. Finally, we demonstrate the robustness of our findings across alternative persuasion-game settings, highlighting the broader potential of using LLM-generated data to model human decision-making.
研究动机与目标
- 探究 LLM 生成的数据是否可以替代人类选择数据,用于训练预测经济情境下人类决策的模型。
- 评估使用基于 LLM 的代理作为语言型说服博弈中合成参与者的可行性和有效性。
- 确定仅基于 LLM 生成数据训练的模型是否能够超越基于真实人类数据训练的模型,在预测人类行为方面表现更优。
- 分析多样化 LLM 人格对合成训练数据质量与多样性的影响。
- 探索在何种条件下,基于 LLM 的合成数据会成为人类数据的更优替代方案,用于人类选择预测。
提出的方法
- 本研究采用 Apel 等人(2022)提出的基于语言的说服博弈框架,其中发送方(专家)通过自然语言信息尝试说服人类决策者(DM)接受酒店交易。
- 使用基于 LLM 的代理来模拟发送方和接收方角色,生成无需任何人类选择数据的合成互动数据。
- 采用多种人格类型(例如:中立、过度积极、怀疑)以多样化合成数据,提升模型泛化能力。
- 仅使用 LLM 生成的数据训练预测模型,以预测人类 DM 对发送方信息的响应。
- 应用 Shapley 值量化每种人格类型对模型整体预测性能的边际贡献。
- 通过与基于真实人类数据训练的模型及语言基线模型的对比,评估该方法的性能。
实验结果
研究问题
- RQ1仅基于 LLM 生成数据训练的模型,是否能够以与基于真实人类数据训练的模型相当或更高的准确度,预测语言型说服博弈中的人类选择?
- RQ2合成数据生成中 LLM 人格的多样性如何影响最终模型的预测性能?
- RQ3在哪些专家策略下,LLM 生成的数据方法优于或劣于人类数据?
- RQ4每种人格类型对合成数据集预测质量的边际贡献是什么?
- RQ5基于 LLM 的合成数据是否能够减少在经济选择预测中对昂贵且涉及隐私敏感的人类数据收集的需求?
主要发现
- 当 LLM 生成的数据集足够大时,仅基于 LLM 生成数据训练的模型在预测准确率上优于基于真实人类选择数据训练的模型。
- 即使在简单专家策略(SendBest)下,即专家始终发送最佳可能的评价(无论酒店质量如何),基于 LLM 的方法仍能实现更优的预测准确率。
- 在数据生成中使用多样化 LLM 人格的混合组合,相比仅使用默认(中立)人格,可显著降低达到特定准确率水平所需的样本量。
- Shapley 值分析显示,所有人格类型对合成数据集预测能力的贡献几乎均匀,表明其贡献均衡且不可或缺。
- 基于 LLM 的合成数据方法始终优于语言基线模型,表明上下文感知的代理行为对准确建模人类行为至关重要。
- 该方法显著减少了数据收集过程中对人类参与者的依赖,提升了训练效率与可扩展性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。