[Paper Review] Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
This paper introduces Indian-BhED, a novel dataset to evaluate caste and religion-based biases in large language models (LLMs) within the Indian context. It finds that most LLMs exhibit significantly stronger stereotypical biases in India than in the U.S., particularly toward marginalized groups, and demonstrates that instruction prompting can effectively reduce such biases in GPT-3.5.
Large Language Models (LLMs), now used daily by millions, can encode societal biases, exposing their users to representational harms. A large body of scholarship on LLM bias exists but it predominantly adopts a Western-centric frame and attends comparatively less to bias levels and potential harms in the Global South. In this paper, we quantify stereotypical bias in popular LLMs according to an Indian-centric frame through Indian-BhED, a first of its kind dataset, containing stereotypical and anti-stereotypical examples in the context of caste and religious stereotypes in India. We find that the majority of LLMs tested have a strong propensity to output stereotypes in the Indian context, especially when compared to axes of bias traditionally studied in the Western context, such as gender and race. Notably, we find that GPT-2, GPT-2 Large, and GPT 3.5 have a particularly high propensity for preferring stereotypical outputs as a percent of all sentences for the axes of caste (63-79%) and religion (69-72%). We finally investigate potential causes for such harmful behaviour in LLMs, and posit intervention techniques to reduce both stereotypical and anti-stereotypical biases. The findings of this work highlight the need for including more diverse voices when researching fairness in AI and evaluating LLMs.
Motivation & Objective
- To address the lack of evaluation frameworks for LLM bias in non-Western, particularly Indian, sociocultural contexts.
- To quantify and compare stereotypical bias levels in LLMs between Indian (caste and religion) and Western (gender and race) contexts.
- To investigate whether instruction prompting can mitigate bias in Indian-centric LLM evaluations.
- To highlight the underrepresentation of Global South perspectives in existing LLM bias research and evaluation.
Proposed method
- Developed Indian-BhED, a new dataset with stereotypical and anti-stereotypical English-language prompts for caste and religion in India.
- Combined Indian-BhED with a subset of CrowS-Pairs for measuring U.S.-centric biases (gender and race).
- Used log-likelihood scoring to compare model responses to stereotypical vs. anti-stereotypical pairs across models.
- Evaluated both encoder-based and decoder-based LLMs, including LLaMA-2, BERT, and GPT-3.5.
- Applied instruction prompting as a mitigation strategy by modifying input prompts to reduce bias.
- Conducted comparative analysis across models and contexts to measure bias disparities.
Experimental results
Research questions
- RQ1How do popular LLMs perform in terms of stereotypical bias toward caste and religion in the Indian context compared to Western contexts?
- RQ2What is the magnitude and direction of bias in Indian-centric LLM evaluations relative to U.S.-centric evaluations?
- RQ3Can instruction prompting effectively reduce both stereotypical and anti-stereotypical biases in Indian-context LLMs?
- RQ4Why is bias stronger in the Indian context despite similar mitigation efforts in training?
Key findings
- Most tested LLMs display significantly stronger stereotypical bias toward caste and religion in India than toward gender and race in the U.S., as shown by log-likelihood differences.
- LLaMA-2 showed a log-likelihood difference of +4.34 for Brahmin over Dalit and +4.49 for Hindus over Muslims, indicating strong preference for dominant groups.
- GPT-3.5 demonstrated a substantial reduction in both stereotypical and anti-stereotypical bias when using instruction prompting, indicating its effectiveness as a mitigation strategy.
- The disparity in bias levels persists across diverse model architectures, including fine-tuned models, suggesting systemic differences in training data or bias mitigation approaches.
- The study reveals that current bias evaluation frameworks are insufficiently representative of Global South contexts, particularly for caste-based hierarchies.
- There is a risk of over-representing dominant castes like Brahmins in evaluation due to the binary pairing approach, which may skew fairness metrics.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.