[Paper Review] Diminished Diversity-of-Thought in a Standard Large Language Model
The paper tests GPT-3.5 as a proxy for human participants in social science replications and documents a phenomenon where certain prompts yield near-zero variation in answers, challenging the validity of LLMs as general replacements for human subjects.
We test whether Large Language Models (LLMs) can be used to simulate human participants in social-science studies. To do this, we run replications of 14 studies from the Many Labs 2 replication project with OpenAI's text-davinci-003 model, colloquially known as GPT3.5. Based on our pre-registered analyses, we find that among the eight studies we could analyse, our GPT sample replicated 37.5% of the original results and 37.5% of the Many Labs 2 results. However, we were unable to analyse the remaining six studies due to an unexpected phenomenon we call the "correct answer" effect. Different runs of GPT3.5 answered nuanced questions probing political orientation, economic preference, judgement, and moral philosophy with zero or near-zero variation in responses: with the supposedly "correct answer." In one exploratory follow-up study, we found that a "correct answer" was robust to changing the demographic details that precede the prompt. In another, we found that most but not all "correct answers" were robust to changing the order of answer choices. One of our most striking findings occurred in our replication of the Moral Foundations Theory survey results, where we found GPT3.5 identifying as a political conservative in 99.6% of the cases, and as a liberal in 99.3% of the cases in the reverse-order condition. However, both self-reported 'GPT conservatives' and 'GPT liberals' showed right-leaning moral foundations. Our results cast doubts on the validity of using LLMs as a general replacement for human participants in the social sciences. Our results also raise concerns that a hypothetical AI-led future may be subject to a diminished diversity-of-thought.
Motivation & Objective
- Assess whether a standard LLM (GPT-3.5) can simulate human participants in social-science replication studies.
- Measure replication success relative to the Many Labs 2 project across multiple tasks.
- Identify and characterize phenomena that undermine diversity of thought in LLM responses.
Proposed method
- Replicate 14 Many Labs 2 studies using OpenAI GPT-3.5 (text-davinci-003).
- Pre-registered analyses to compare GPT outputs with original and Many Labs 2 results.
- Analyze eight analysable studies and report replication rate; document unexpected near-zero variation in responses.
- Conduct exploratory follow-ups varying demographic details and prompt order to test robustness of responses.
- Examine Moral Foundations Theory survey outcomes under different prompt conditions.
Experimental results
Research questions
- RQ1Can GPT-3.5 replicate a substantial portion of the original Many Labs 2 results?
- RQ2Do GPT-3.5 responses exhibit sufficient diversity to be considered valid surrogates for human participants?
- RQ3What phenomena (e.g., 'correct answer' effect) emerge that limit the use of LLMs for social science replications?
- RQ4How robust are GPT-3.5-derived conclusions to changes in demographics and answer-order prompts?
- RQ5What are the implications for AI-led future scenarios concerning diversity of thought?
Key findings
- GPT-3.5 replicated 37.5% of the original results across eight analysable studies.
- GPT-3.5 replicated 37.5% of the Many Labs 2 results.
- An unexpected 'correct answer' effect produced zero or near-zero variation in responses to nuanced questions.
- In exploratory follow-ups, the 'correct answer' was robust to demographic variations preceding the prompt.
- In Moral Foundations Theory replications, GPT-3.5 identified as conservative in 99.6% of reverse-order cases and liberal in 99.3% of cases, yet both groups showed right-leaning moral foundations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.