[Paper Review] Generative Agent Simulations of 1,000 People
The paper builds generative agents from self-report data (interviews and surveys) to simulate individuals, achieving high accuracy and reducing bias across groups without task-specific training data.
The promise of human behavioral simulation--general-purpose computational agents that replicate human behavior across domains--could enable broad applications in policymaking and social science. We present a novel agent architecture that simulates the attitudes and behaviors of 1,052 real individuals--applying large language models to qualitative interviews about their lives, then measuring how well these agents replicate the attitudes and behaviors of the individuals that they represent. The generative agents replicate participants' responses on the General Social Survey 85% as accurately as participants replicate their own answers two weeks later, and perform comparably in predicting personality traits and outcomes in experimental replications. Our architecture reduces accuracy biases across racial and ideological groups compared to agents given demographic descriptions. This work provides a foundation for new tools that can help investigate individual and collective behavior.
Motivation & Objective
- Demonstrate that LLM-based agents grounded in self-report data can generalize across diverse outcomes without task-specific training.
- Evaluate two data sources—semi-structured interviews and structured surveys—for agent accuracy.
- Assess whether combining sources improves performance and reduces disparities.
- Show whether agents can predict personality traits and behaviors in experiments similarly to humans.
Proposed method
- Construct agents from two-hour semi-structured interviews (American Voices Project) and/or structured surveys (General Social Survey, Big Five).
- Ground LLMs with self-report data to simulate individuals rather than using demographics alone.
- Evaluate agent accuracy on held-out GSS items against two-week test-retest consistency.
- Compare performance of interview-only, surveys-only, and combined data on prediction tasks.
- Assess racial and ideological group disparities in accuracy relative to demographics-only baselines.
Experimental results
Research questions
- RQ1Can LLM agents grounded in self-reports generalize to multiple outcomes without task-specific training data?
- RQ2How do different data sources (interviews, surveys, or both) affect agent accuracy and consistency?
- RQ3Do self-report-grounded agents reduce racial and ideological disparities in predictive accuracy?
- RQ4How do agent predictions compare to human benchmarks on personality traits and behaviors?
- RQ5What is the impact of combining interview and survey data on simulation quality?
Key findings
- Interview-only agents achieve 83% accuracy relative to participants' two-week test-retest; surveys-only agents achieve 82%; combined data agents achieve 86%.
- Baseline agents prompted only with demographics achieve 74% accuracy.
- Agents predict personality traits and behaviors with similar accuracy to humans on experiments.
- Self-report-grounded agents reduce disparities in accuracy across racial and ideological groups compared with demographics-only baselines.
- Using rich self-report data enables general-purpose simulation across outcomes without task-specific training data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.