[Paper Review] Learning to be Homo Economicus: Can an LLM Learn Preferences from Choice
This paper investigates whether GPT can learn economic preferences from choice data and provide personalized recommendations. By prompting GPT to act as a decision-maker or recommendation system, the study finds that GPT consistently aligns with expected utility maximization and can personalize recommendations based on risk aversion, though it shows limited ability to capture disappointment aversion, demonstrating partial but meaningful learning from data.
This paper explores the use of Large Language Models (LLMs) as decision aids, with a focus on their ability to learn preferences and provide personalized recommendations. To establish a baseline, we replicate standard economic experiments on choice under risk (Choi et al., 2007) with GPT, one of the most prominent LLMs, prompted to respond as (i) a human decision maker or (ii) a recommendation system for customers. With these baselines established, GPT is provided with a sample set of choices and prompted to make recommendations based on the provided data. From the data generated by GPT, we identify its (revealed) preferences and explore its ability to learn from data. Our analysis yields three results. First, GPT's choices are consistent with (expected) utility maximization theory. Second, GPT can align its recommendations with people's risk aversion, by recommending less risky portfolios to more risk-averse decision makers, highlighting GPT's potential as a personalized decision aid. Third, however, GPT demonstrates limited alignment when it comes to disappointment aversion.
Motivation & Objective
- To evaluate GPT’s ability to learn preferences from choice data and serve as a personalized decision aid.
- To test whether GPT’s behavior is consistent with revealed preference theory and expected utility maximization.
- To assess the quality of GPT’s personalized recommendations as a function of sample size.
- To compare GPT’s performance as a decision-maker versus a recommendation system.
- To develop a methodological framework for evaluating LLMs in economic decision-making contexts.
Proposed method
- Replicating Choi et al. (2007) choice experiments under risk, prompting GPT to respond as a human decision-maker or as a recommendation system.
- Using revealed preference techniques to test consistency with utility maximization and expected utility maximization.
- Parametrically recovering preference parameters from GPT’s responses using methods from Afriat (1967) and Echenique et al. (2023).
- Providing GPT with simulated choice data from human subjects to assess its ability to infer preferences and make personalized recommendations.
- Varying the size of the provided choice dataset (from 1 to 175 samples) to measure the accuracy-size trade-off in personalization.
- Comparing GPT’s recommendations to true human risk aversion parameters via scatter plots and regression analysis.
Experimental results
Research questions
- RQ1Can GPT’s choices be consistent with expected utility maximization theory when prompted as a decision-maker?
- RQ2To what extent can GPT learn and reflect individual risk aversion from a small set of choice data when acting as a recommendation system?
- RQ3Does GPT’s ability to personalize recommendations improve with increasing sample size of choice data?
- RQ4How well does GPT’s recommendation behavior align with human preferences, particularly in capturing risk aversion versus disappointment aversion?
- RQ5Can a simple prompt enable GPT to learn economic preferences without fine-tuning or explicit instruction on preference structures?
Key findings
- GPT’s choices as a decision-maker are consistent with expected utility maximization, as confirmed by revealed preference analysis.
- When acting as a recommendation system, GPT successfully personalizes portfolio recommendations based on inferred risk aversion, with performance improving significantly as sample size increases.
- With only one choice sample, GPT’s recommendations show no meaningful personalization (regression coefficient = 0.002, p = 0.960), but performance improves markedly with larger datasets.
- The slope of the regression line between true human risk aversion and GPT-estimated risk aversion increases with sample size, indicating strong learning from data.
- GPT fails to adequately capture disappointment aversion, suggesting limitations in modeling higher-order risk preferences.
- GPT-4 outperforms GPT-3.5-turbo in stability and personalization, indicating that model size and architecture significantly influence performance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.