Skip to main content
QUICK REVIEW

[论文解读] The Emergence of Economic Rationality of GPT

Yi‐Ting Chen, Tracy Xiao Liu|arXiv (Cornell University)|May 22, 2023
Topic Modeling被引用 5
一句话总结

本研究通过揭示偏好理论,评估GPT在风险、时间、社会偏好和食物四个领域中的预算分配决策,检验其经济理性。与人类受试者相比,GPT在类似实验中表现出更高的理性得分,其一致性对人口统计和随机性变化具有鲁棒性,但对框架设计和选择格式敏感。

ABSTRACT

As large language models (LLMs) like GPT become increasingly prevalent, it is essential that we assess their capabilities beyond language processing. This paper examines the economic rationality of GPT by instructing it to make budgetary decisions in four domains: risk, time, social, and food preferences. We measure economic rationality by assessing the consistency of GPT's decisions with utility maximization in classic revealed preference theory. We find that GPT's decisions are largely rational in each domain and demonstrate higher rationality score than those of human subjects in a parallel experiment and in the literature. Moreover, the estimated preference parameters of GPT are slightly different from human subjects and exhibit a lower degree of heterogeneity. We also find that the rationality scores are robust to the degree of randomness and demographic settings such as age and gender, but are sensitive to contexts based on the language frames of the choice situations. These results suggest the potential of LLMs to make good decisions and the need to further understand their capabilities, limitations, and underlying mechanisms.

研究动机与目标

  • 评估GPT在多样化领域中的决策是否表现出经济理性。
  • 利用揭示偏好理论,比较GPT与人类受试者的理性水平。
  • 检验GPT理性对随机性、人口统计框架和选择情境变化的鲁棒性。
  • 探究大语言模型与人类在潜在机制及偏好参数异质性方面的差异。
  • 评估大语言模型决策质量对人工智能在未来经济与社会应用中影响的启示。

提出的方法

  • GPT被提示在每个领域内进行25次预算分配决策,将100分在两种商品之间分配,商品价格各不相同。
  • 理性通过广义 revealed preference 公理(GARP)进行衡量,该公理是预算约束下效用最大化的必要且充分条件。
  • 构建了四种不同的决策环境:风险资产、跨期选择、社会转移和食物偏好。
  • 理性得分通过 Afriat 效率指数(CCEI)以及对数价格与对数数量比率之间的斯皮尔曼等级相关系数计算得出。
  • 实验包括价格框架的变化(如成本 vs. 收益)、选择格式(连续 vs. 离散)以及人口统计框架(年龄、性别、种族)。
  • 通过347名美国参与者的平行人类受试者实验,实现了GPT与人类理性得分的直接比较。

实验结果

研究问题

  • RQ1在风险、时间、社会偏好和食物偏好等预算分配任务中,GPT在多大程度上表现出经济理性?
  • RQ2在相同实验设置下,GPT的理性得分与人类受试者相比如何?
  • RQ3GPT的理性是否对随机性、人口统计框架和选择情境的变化具有鲁棒性?
  • RQ4GPT估计的偏好参数在异质性和一致性方面与人类相比如何?
  • RQ5语言框架和选择格式在多大程度上调节了GPT的理性得分?

主要发现

  • 在本实验及更广泛文献范围内,GPT的理性得分显著高于人类受试者。
  • GPT的理性对人口统计设置(年龄、性别、种族、教育程度)以及响应中随机性的变化具有鲁棒性。
  • 当选择任务采用不同语言框架(如成本 vs. 收益框架)或以离散选择格式呈现时,理性得分显著下降。
  • GPT估计的偏好参数异质性低于人类受试者,表明其决策模式更具一致性。
  • GPT选择中对数价格与对数数量比率之间的相关性高于人类受试者,表明其与效用最大化的对齐程度更高。
  • 在风险、时间、社会和食物四个领域中,GPT的决策始终与GARP一致,证实其揭示偏好的一致性水平极高。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。