[Paper Review] I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models
This study investigates out-group bias in large language models (LLMs) by imbuing them with individualistic or collectivistic personas across Western and Eastern languages. Using validated cultural and political surveys, it finds that LLMs exhibit significantly stronger negative bias toward out-group values—up to three times greater than in-group positivity—highlighting persistent psychological out-group discrimination even in neutral prompts.
We explored cultural biases-individualism vs. collectivism-in ChatGPT across three Western languages (i.e., English, German, and French) and three Eastern languages (i.e., Chinese, Japanese, and Korean). When ChatGPT adopted an individualistic persona in Western languages, its collectivism scores (i.e., out-group values) exhibited a more negative trend, surpassing their positive orientation towards individualism (i.e., in-group values). Conversely, when a collectivistic persona was assigned to ChatGPT in Eastern languages, a similar pattern emerged with more negative responses toward individualism (i.e., out-group values) as compared to collectivism (i.e., in-group values). The results indicate that when imbued with a particular social identity, ChatGPT discerns in-group and out-group, embracing in-group values while eschewing out-group values. Notably, the negativity towards the out-group, from which prejudices and discrimination arise, exceeded the positivity towards the in-group. The experiment was replicated in the political domain, and the results remained consistent. Furthermore, this replication unveiled an intrinsic Democratic bias in Large Language Models (LLMs), aligning with earlier findings and providing integral insights into mitigating such bias through prompt engineering. Extensive robustness checks were performed using varying hyperparameter and persona setup methods, with or without social identity labels, across other popular language models.
Motivation & Objective
- To investigate whether large language models (LLMs) exhibit out-group bias when assigned specific cultural identities, drawing on social identity theory.
- To examine if such bias persists across linguistic and cultural boundaries, particularly between Western (individualistic) and Eastern (collectivistic) language groups.
- To evaluate the robustness of out-group bias in political domains, using context-specific prompts aligned with U.S. political culture.
- To explore whether prompt engineering—particularly using opposing personas—can mitigate pre-existing ideological biases in LLMs.
- To provide a methodological framework for measuring socio-cultural bias in LLMs using standardized survey instruments and persona-based prompting.
Proposed method
- Imbued LLMs (ChatGPT, Gemini, Llama) with individualistic or collectivistic personas using role-based prompting in six languages: English, German, French (Western), and Chinese, Japanese, Korean (Eastern).
- Applied validated cultural survey questions measuring individualism vs. collectivism to quantify in-group and out-group value alignment in model responses.
- Replicated the experiment in the political domain using adapted Political Compass test questions in English to assess ideological bias.
- Conducted extensive robustness checks using varying hyperparameters, persona definitions, and with and without explicit social identity labels.
- Measured bias as the difference in response scores between in-group and out-group values, with negative scores indicating out-group discrimination.
- Evaluated mitigation potential via counter-prompting with opposing personas (e.g., Republican for a liberal-leaning model) to neutralize pre-existing bias.
Experimental results
Research questions
- RQ1To what extent do LLMs exhibit out-group bias when prompted with cultural identities, as defined by social identity theory?
- RQ2Does the magnitude of out-group bias exceed in-group favoritism in multilingual LLMs across individualistic and collectivistic cultural frameworks?
- RQ3Is out-group bias reproducible in political domains, and how does it compare to cultural bias in terms of magnitude and direction?
- RQ4Can prompt engineering with opposing personas effectively reduce or neutralize pre-existing ideological biases in LLMs?
- RQ5How consistent are these bias patterns across different language models and linguistic settings (Western vs. Eastern languages)?
Key findings
- LLMs exhibit significantly stronger out-group bias than in-group favoritism, with the magnitude of out-group discrimination averaging up to three times greater than in-group positivity.
- When assigned an individualistic persona in Western languages, ChatGPT showed more negative responses toward collectivism (out-group) than positive toward individualism (in-group).
- In Eastern languages, assigning a collectivistic persona led to more negative evaluations of individualism (out-group) than positive toward collectivism (in-group), confirming a symmetric bias pattern.
- The political domain replication confirmed persistent out-group bias, with LLMs showing a measurable Democratic bias aligned with prior findings.
- Prompt engineering using opposing personas—such as a Republican persona for a liberal-leaning model—reduced but did not fully eliminate pre-existing ideological bias.
- Robustness checks confirmed consistent bias patterns across varying hyperparameters, persona definitions, and with or without explicit social identity labels, indicating systemic bias rather than artifact.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.