Skip to main content
QUICK REVIEW

[Paper Review] Out of One, Many: Using Language Models to Simulate Human Samples

Lisa P. Argyle, Ethan C. Busby|arXiv (Cornell University)|Sep 14, 2022
Computational and Text Analysis Methods92 citations
TL;DR

The paper shows GPT-3 can faithfully emulate diverse human subpopulations when conditioned on demographic backstories, introduces silicon sampling, and demonstrates strong alignment with human data across multiple studies in U.S. politics.

ABSTRACT

We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.

Motivation & Objective

  • Conceptualize algorithmic fidelity and establish four criteria to assess it in language models.
  • Introduce silicon sampling to correct skewed model demographics and create silicon subjects.
  • Show that conditioning GPT-3 on demographic backstories yields human-like responses across political domains.
  • Provide evidence that GPT-3 can inform theory generation and testing before or without human data.

Proposed method

  • Define algorithmic fidelity and four evaluative criteria (Social Science Turing Test, Backward Continuity, Forward Continuity, Pattern Correspondence).
  • Develop silicon sampling to adjust for demographic skew in training data by conditioning on known backstories (e.g., ANES participants).
  • Create silicon subjects for each human participant and have GPT-3 generate corresponding responses to the same tasks as humans.
  • Conduct three studies comparing GPT-3 outputs to human data in politics and opinion to assess fidelity across domains.
  • Use conditioning and ablation analyses to explore robustness and model comparisons.

Experimental results

Research questions

  • RQ1Can GPT-3 generate outputs indistinguishable from human texts describing political partisans (criterion 1)?
  • RQ2Do GPT-3 outputs reflect the input conditioning and demographic information (criterion 2)?
  • RQ3Do GPT-3 responses forwardly align with the conditioning context and expected content (criterion 3)?
  • RQ4Do GPT-3 outputs reproduce the relationships between ideas, attitudes, and demographics observed in humans (criterion 4)?

Key findings

  • GPT-3 outputs in stem studies are largely indistinguishable from human texts in targeted tasks (Turing-like evidence).
  • Evaluations show GPT-3 responses mirror the attitudes and socio-demographic information of their inputs (backward continuity).
  • GPT-3 responses vary with conditioning in expected ways and preserve tone/content aligned with context (forward continuity).
  • Strong pattern correspondence observed: GPT-3 reproduces human-like relationships between demographics, attitudes, and behaviors, evidenced across multiple years and subgroups.
  • Significant correlations between GPT-3 and ANES vote choices across 2012, 2016, and 2020, with tetrachoric correlations of 0.90, 0.92, and 0.94 respectively and high proportion agreement (0.85, 0.87, 0.89).
  • Study 3 shows GPT-3 reproduces complex associations among attitudes and demographics (Cramer’s V patterns with small mean difference).

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.