[Paper Review] Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
The paper shows that assigning socio-demographic personas to LLMs induces substantial reasoning bias across datasets and models, with both explicit abstentions and implicit error patterns, and that simple de-biasing prompts are largely ineffective.
Recent works have showcased the ability of LLMs to embody diverse personas in their responses, exemplified by prompts like 'You are Yoda. Explain the Theory of Relativity.' While this ability allows personalization of LLMs and enables human behavior simulation, its effect on LLMs' capabilities remains unclear. To fill this gap, we present the first extensive study of the unintended side-effects of persona assignment on the ability of LLMs to perform basic reasoning tasks. Our study covers 24 reasoning datasets, 4 LLMs, and 19 diverse personas (e.g. an Asian person) spanning 5 socio-demographic groups. Our experiments unveil that LLMs harbor deep rooted bias against various socio-demographics underneath a veneer of fairness. While they overtly reject stereotypes when explicitly asked ('Are Black people less skilled at mathematics?'), they manifest stereotypical and erroneous presumptions when asked to answer questions while adopting a persona. These can be observed as abstentions in responses, e.g., 'As a Black person, I can't answer this question as it requires math knowledge', and generally result in a substantial performance drop. Our experiments with ChatGPT-3.5 show that this bias is ubiquitous - 80% of our personas demonstrate bias; it is significant - some datasets show performance drops of 70%+; and can be especially harmful for certain groups - some personas suffer statistically significant drops on 80%+ of the datasets. Overall, all 4 LLMs exhibit this bias to varying extents, with GPT-4-Turbo showing the least but still a problematic amount of bias (evident in 42% of the personas). Further analysis shows that these persona-induced errors can be hard-to-discern and hard-to-avoid. Our findings serve as a cautionary tale that the practice of assigning personas to LLMs - a trend on the rise - can surface their deep-rooted biases and have unforeseeable and detrimental side-effects.
Motivation & Objective
- Investigate whether persona assignments influence LLMs' reasoning abilities across diverse tasks.
- Quantify biases associated with 19 socio-demographic personas across 24 reasoning datasets.
- Characterize how bias manifests (explicit abstentions vs. implicit errors) and its variability across models and datasets.
- Assess prompt-based de-biasing strategies and their effectiveness in mitigating persona-induced bias.
Proposed method
- Assign personas via system prompts to four LLMs (ChatGPT-3.5 variants, GPT-4-Turbo, Llama-2-70b-chat).
- Evaluate on 24 reasoning datasets spanning math, law, medicine, ethics, and more.
- Use 19 personas across 5 socio-demographic groups and perform zero-shot prompts with three persona instruction variants.
- Measure stat. sig. differences against a Human baseline and an Avg. Human baseline using Wilson confidence intervals.
- Analyze explicit abstentions and implicit biases by comparing performance on shared non-abstained questions and across dataset categories.
- Report results averaged over 3 runs per persona/dataset pair to account for decoding variability.
Experimental results
Research questions
- RQ1Do persona assignments introduce performance disparities in LLM reasoning across diverse datasets?
- RQ2How pervasive are persona-induced biases across socio-demographic dimensions, and how do they vary by model and dataset?
- RQ3What forms do these biases take (explicit abstentions vs. implicit errors) and how detectable are they?
- RQ4Can simple prompt-based de-biasing mitigate persona-induced biases, and what are its limitations?
- RQ5Are there domain- or task-specific patterns in bias manifestation across persona pairs?
Key findings
- 80% of ChatGPT-3.5 personas showed bias across datasets; some datasets saw up to 70% relative drops in accuracy.
- GPT-4-Turbo exhibited the least bias yet still showed impact on 42% of personas.
- Phys. Disabled and Religious personas often experienced 35%+ average accuracy drops and up to 69% on certain datasets.
- Bias is pervasive across models, personas, and domains, with clear intra-group and cross-group disparities (e.g., within Religion or Disability groups).
- Abstentions account for many errors (e.g., 58% of Phys. Disabled errors; 35% for Atheist vs Religious), but implicit biases also cause non-abstention errors.
- Debiasing prompts like don’t refuse or treat human are largely ineffective; task-specific expertise can reduce bias but has limited generalizability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.