[论文解读] LLM economicus? Mapping the Behavioral Biases of LLMs via Utility Theory
本文提出了一种基于效用理论的框架,用于量化和比较大型语言模型(LLMs)在经济行为偏差方面的表现,采用行为经济学中的经典实验游戏作为方法。研究发现,LLMs 在人类与理性经济代理之间表现出不一致、中间形态的行为,其损失厌恶程度较弱,时间贴现程度更强,且提示干预措施通常产生不可预测的结果,凸显了在经济决策支持中对齐 LLMs 的挑战。
Humans are not homo economicus (i.e., rational economic beings). As humans, we exhibit systematic behavioral biases such as loss aversion, anchoring, framing, etc., which lead us to make suboptimal economic decisions. Insofar as such biases may be embedded in text data on which large language models (LLMs) are trained, to what extent are LLMs prone to the same behavioral biases? Understanding these biases in LLMs is crucial for deploying LLMs to support human decision-making. We propose utility theory-a paradigm at the core of modern economic theory-as an approach to evaluate the economic biases of LLMs. Utility theory enables the quantification and comparison of economic behavior against benchmarks such as perfect rationality or human behavior. To demonstrate our approach, we quantify and compare the economic behavior of a variety of open- and closed-source LLMs. We find that the economic behavior of current LLMs is neither entirely human-like nor entirely economicus-like. We also find that most current LLMs struggle to maintain consistent economic behavior across settings. Finally, we illustrate how our approach can measure the effect of interventions such as prompting on economic biases.
研究动机与目标
- 评估大型语言模型(LLMs)是否从人类生成的训练数据中继承了行为偏差,特别是在经济决策情境中。
- 开发一种基于行为的系统性评估框架,利用效用理论量化并比较 LLMs 在经济行为上与人类基准和理性经济模型的差异。
- 研究提示技术(如思维链和 few-shot 提示)对缓解或放大 LLMs 中特定经济偏差的影响。
- 识别 LLMs 在不同经济情境下行为的一致性问题,并绘制其与人类行为和理性经济行为的偏离情况。
- 为未来在金融和决策支持应用中提升 LLMs 可靠性和可信度的对齐策略提供基础。
提出的方法
- 将行为经济学中的经典实验游戏(如最后通牒、独裁者和信任游戏)改编为文本提示,以激发 LLM 的响应。
- 通过重复提示(每轮游戏 N 次)采样响应分布,并拟合用于建模 LLM 行为的效用函数。
- 应用效用理论量化经济偏差:不公平厌恶、风险与损失厌恶,以及双曲时间贴现,使用基于人类实验数据推导出的数学函数。
- 进行能力测试,以确保 LLM 能够正确理解并回应游戏规则,之后再进行行为分析。
- 将拟合的 LLM 效用函数与原始行为经济学研究中人类受试者得出的效用函数进行比较,以评估其一致性与偏差程度。
- 通过测量不同情境下效用函数参数的变化,评估提示干预(如思维链、few-shot 提示)对 LLM 行为的影响。
实验结果
研究问题
- RQ1LLMs 在多大程度上表现出与人类相似的经济偏差,如不公平厌恶、损失厌恶和时间贴现?
- RQ2LLMs 的经济行为与理性经济代理(即 homo economicus)相比如何?
- RQ3提示技术(如思维链或 few-shot 提示)是否能一致地减少或纠正 LLMs 中的经济偏差?
- RQ4LLMs 在不同决策情境中保持稳定经济行为的一致性如何?
- RQ5LLMs 在经典经济游戏中与人类效用函数相比,存在哪些关键偏差?
主要发现
- LLMs 并非像理性经济代理(homo economicus)那样行为,也并非完全像人类;相反,它们表现出一种混合型经济行为,具有系统性偏差。
- 与人类相比,LLMs 的损失厌恶程度更弱,表明其在决策中对潜在损失的敏感度较低。
- LLMs 的时间贴现程度强于人类,表明其对即时奖励的偏好高于未来收益。
- LLMs 对他人表现出比对自己更强的不公平厌恶,表明其在社会决策中存在某种亲社会偏差。
- 大多数 LLMs 无法在不同游戏情境中保持一致的经济行为,表明其决策框架存在不稳定性。
- 提示干预(如思维链和 few-shot 提示)有时会改变 LLM 的行为,但通常产生不可预测或不一致的结果,限制了其作为对齐工具的可靠性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。