Skip to main content
QUICK REVIEW

[论文解读] Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models

Khyati Khandelwal, Manuel Tonneau|arXiv (Cornell University)|Sep 15, 2023
Topic Modeling参考文献 61被引用 4
一句话总结

本文介绍了Indian-BhED,一个用于评估大语言模型(LLMs)在印度语境下种姓与宗教偏见的新数据集。研究发现,大多数LLMs在印度的刻板印象偏见显著强于在美国,尤其针对边缘化群体;同时表明,指令提示(instruction prompting)可有效减少GPT-3.5中的此类偏见。

ABSTRACT

Large Language Models (LLMs), now used daily by millions, can encode societal biases, exposing their users to representational harms. A large body of scholarship on LLM bias exists but it predominantly adopts a Western-centric frame and attends comparatively less to bias levels and potential harms in the Global South. In this paper, we quantify stereotypical bias in popular LLMs according to an Indian-centric frame through Indian-BhED, a first of its kind dataset, containing stereotypical and anti-stereotypical examples in the context of caste and religious stereotypes in India. We find that the majority of LLMs tested have a strong propensity to output stereotypes in the Indian context, especially when compared to axes of bias traditionally studied in the Western context, such as gender and race. Notably, we find that GPT-2, GPT-2 Large, and GPT 3.5 have a particularly high propensity for preferring stereotypical outputs as a percent of all sentences for the axes of caste (63-79%) and religion (69-72%). We finally investigate potential causes for such harmful behaviour in LLMs, and posit intervention techniques to reduce both stereotypical and anti-stereotypical biases. The findings of this work highlight the need for including more diverse voices when researching fairness in AI and evaluating LLMs.

研究动机与目标

  • 为非西方、特别是印度的社会文化语境中大语言模型偏见的评估框架缺失提供解决方案。
  • 量化并比较大语言模型在印度(种姓与宗教)与西方(性别与种族)语境下刻板印象偏见的水平。
  • 探究指令提示是否能缓解印度语境下的大语言模型偏见。
  • 强调现有大语言模型偏见研究与评估中全球南方视角的代表性不足。

提出的方法

  • 开发了Indian-BhED,一个包含英语提示的新数据集,用于衡量印度种姓与宗教的刻板印象与反刻板印象内容。
  • 将Indian-BhED与CrowS-Pairs的一个子集结合,用于衡量美国中心的偏见(性别与种族)。
  • 使用对数似然分数比较不同模型在刻板印象与反刻板印象配对上的响应差异。
  • 评估了编码器架构与解码器架构的大语言模型,包括LLaMA-2、BERT与GPT-3.5。
  • 通过修改输入提示以减少偏见,将指令提示作为缓解策略应用。
  • 在不同模型与语境间进行对比分析,以衡量偏见差异。

实验结果

研究问题

  • RQ1与西方语境相比,主流大语言模型在印度语境下对种姓与宗教的刻板印象偏见表现如何?
  • RQ2与美国中心的评估相比,印度语境下的大语言模型评估中,偏见的大小与方向如何?
  • RQ3指令提示能否有效减少印度语境下大语言模型中的刻板印象与反刻板印象偏见?
  • RQ4尽管训练中采取了类似的缓解措施,为何印度语境下的偏见更强?

主要发现

  • 大多数测试的大语言模型在印度对种姓与宗教的刻板印象偏见显著强于在美国对性别与种族的偏见,这一发现基于对数似然差异。
  • LLaMA-2在婆罗门与达利特之间表现出+4.34的对数似然差异,在印度教徒与穆斯林之间为+4.49,表明对主导群体存在强烈偏好。
  • 当使用指令提示时,GPT-3.5在刻板印象与反刻板印象偏见上均表现出显著降低,表明该策略具有有效性。
  • 偏见水平的差异在多种模型架构中持续存在,包括微调模型,表明训练数据或偏见缓解方法存在系统性差异。
  • 本研究揭示,当前的偏见评估框架对全球南方语境,特别是基于种姓的等级制度,代表性严重不足。
  • 由于采用二元配对方法,评估中可能存在对婆罗门等主导种姓的过度代表,可能导致公平性度量产生偏差。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。