[Paper Review] LLM economicus? Mapping the Behavioral Biases of LLMs via Utility Theory
This paper proposes a utility theory-based framework to quantify and compare economic behavioral biases in large language models (LLMs), using canonical experimental games from behavioral economics. It finds that LLMs exhibit inconsistent, intermediate behavior between humans and rational economic agents, with weaker loss aversion and stronger time discounting, and that prompting interventions often yield unpredictable results, highlighting challenges in aligning LLMs for economic decision support.
Humans are not homo economicus (i.e., rational economic beings). As humans, we exhibit systematic behavioral biases such as loss aversion, anchoring, framing, etc., which lead us to make suboptimal economic decisions. Insofar as such biases may be embedded in text data on which large language models (LLMs) are trained, to what extent are LLMs prone to the same behavioral biases? Understanding these biases in LLMs is crucial for deploying LLMs to support human decision-making. We propose utility theory-a paradigm at the core of modern economic theory-as an approach to evaluate the economic biases of LLMs. Utility theory enables the quantification and comparison of economic behavior against benchmarks such as perfect rationality or human behavior. To demonstrate our approach, we quantify and compare the economic behavior of a variety of open- and closed-source LLMs. We find that the economic behavior of current LLMs is neither entirely human-like nor entirely economicus-like. We also find that most current LLMs struggle to maintain consistent economic behavior across settings. Finally, we illustrate how our approach can measure the effect of interventions such as prompting on economic biases.
Motivation & Objective
- To assess whether large language models (LLMs) inherit behavioral biases from human-generated training data, particularly in economic decision-making contexts.
- To develop a systematic, behavior-based evaluation framework using utility theory to quantify and compare LLMs' economic behavior against human benchmarks and rational economic models.
- To investigate the impact of prompting techniques—such as chain-of-thought and few-shot prompting—on the mitigation or amplification of specific economic biases in LLMs.
- To identify inconsistencies in LLM behavior across different economic settings and to map deviations from human-like and rational economic behavior.
- To provide a foundation for future alignment strategies that improve the reliability and trustworthiness of LLMs in financial and decision-support applications.
Proposed method
- Adapts classic experimental games from behavioral economics—such as the ultimatum, dictator, and trust games—into text-based prompts to elicit LLM responses.
- Uses repeated prompting (N times per game turn) to sample response distributions and fit utility functions that model LLM behavior.
- Applies utility theory to quantify economic biases: inequity aversion, risk and loss aversion, and hyperbolic time discounting, using mathematical functions derived from human experimental data.
- Performs a competence test to ensure LLMs can correctly interpret and respond to game rules before behavioral analysis.
- Compares fitted LLM utility functions to those derived from human subjects in original behavioral economics studies to assess alignment and deviation.
- Evaluates the effect of prompting interventions (e.g., chain-of-thought, few-shot) on LLM behavior by measuring shifts in utility function parameters across settings.
Experimental results
Research questions
- RQ1To what extent do LLMs exhibit behavioral biases such as inequity aversion, loss aversion, and time discounting similar to humans?
- RQ2How do the economic behaviors of LLMs compare to those of rational economic agents (i.e., homo economicus)?
- RQ3Can prompting techniques like chain-of-thought or few-shot prompting consistently reduce or correct economic biases in LLMs?
- RQ4How consistent are LLMs in maintaining stable economic behavior across different decision-making contexts?
- RQ5What are the key deviations of LLMs from human utility functions in canonical economic games?
Key findings
- LLMs do not behave like rational economic agents (homo economicus) nor fully like humans; instead, they exhibit a hybrid form of economic behavior with systematic deviations.
- LLMs show weaker loss aversion compared to humans, suggesting they are less sensitive to potential losses in decision-making.
- LLMs display stronger time discounting than humans, indicating a greater preference for immediate rewards over future gains.
- LLMs exhibit stronger inequity aversion toward others than toward themselves, suggesting a form of prosocial bias in social decision-making.
- Most LLMs fail to maintain consistent economic behavior across different game settings, indicating instability in their decision-making frameworks.
- Prompting interventions such as chain-of-thought and few-shot prompting sometimes alter LLM behavior but often produce unpredictable or inconsistent results, limiting their reliability as alignment tools.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.