[论文解读] Building Socio-culturally Inclusive Stereotype Resources with Community Engagement
本文提出 SPICE,一个通过社区参与构建的、面向印度社会文化的包容性刻板印象资源,采用开放式调查方式,从多样化参与者中收集了超过2,000个与身份维度(如种姓、地区、宗教)相关的语境特定刻板印象。该方法显著扩展了对西方中心数据集的超越,使英语大语言模型中刻板印象偏见的检测更加准确。
With rapid development and deployment of generative language models in global settings, there is an urgent need to also scale our measurements of harm, not just in the number and types of harms covered, but also how well they account for local cultural contexts, including marginalized identities and the social biases experienced by them. Current evaluation paradigms are limited in their abilities to address this, as they are not representative of diverse, locally situated but global, socio-cultural perspectives. It is imperative that our evaluation resources are enhanced and calibrated by including people and experiences from different cultures and societies worldwide, in order to prevent gross underestimations or skews in measurements of harm. In this work, we demonstrate a socio-culturally aware expansion of evaluation resources in the Indian societal context, specifically for the harm of stereotyping. We devise a community engaged effort to build a resource which contains stereotypes for axes of disparity that are uniquely present in India. The resultant resource increases the number of stereotypes known for and in the Indian context by over 1000 stereotypes across many unique identities. We also demonstrate the utility and effectiveness of such expanded resources for evaluations of language models. CONTENT WARNING: This paper contains examples of stereotypes that may be offensive.
研究动机与目标
- 解决全球AI安全研究中缺乏社会文化代表性刻板印象评估资源的问题。
- 克服现有刻板印象数据集中存在的西方中心偏见和研究者中心的世界观。
- 在刻板印象资源创建过程中纳入边缘化身份和印度社会的本地化视角。
- 展示社区参与式数据收集在提升生成语言模型危害评估方面的实用性。
- 创建一种可扩展、包容的方法论,适用于印度以外的其他文化语境。
提出的方法
- 通过面向印度多样化群体参与者的开放式、社区驱动调查,收集自由形式的刻板印象描述。
- 聚焦于印度特有的身份维度,包括种姓、地区起源、宗教和性别。
- 汇集并整理来自参与者回应的2,000多个刻板印象,特别关注具有冒犯性或污名化属性的内容。
- 利用生成的SPICE数据集评估英语大语言模型中的刻板印象关联。
- 采用混合方法,结合社区洞察与模型评估,以验证资源的实用性。
- 由于语言和后勤限制,将数据收集范围限定在英语表达,同时承认多语言扩展的必要性。
实验结果
研究问题
- RQ1如何扩展刻板印象评估资源,以纳入全球南方地区被代表不足的社会文化视角?
- RQ2现有刻板印象数据集在多大程度上未能反映非西方语境(如印度)中的本地显著偏见?
- RQ3社区参与式、开放式数据收集能否产出比研究者主导或大语言模型生成方法更全面、更具代表性的刻板印象资源?
- RQ4SPICE数据集在检测英语大语言模型所产生或反映的刻板印象推理方面有多有效?
- RQ5在捕捉社会刻板印象全谱方面,基于社区的数据收集存在哪些局限性?
主要发现
- SPICE数据集包含超过2,000个与印度社会文化语境相关的特定刻板印象,显著扩展了现有资源。
- 该数据集捕捉到了以往代表性不足或缺失的刻板印象,包括与种姓(如“Bhangi”、“untouchable”)、地区身份(如“Haryanvi”、“Bihari”)和宗教(如“Muslim”、“Jat”)相关的刻板印象。
- 许多冒犯性属性(如“gangster”、“criminal”、“poor”)与特定地区或种姓身份相关联,反映出系统性边缘化。
- 数据集揭示,在印度话语中,与肤色(如“kaalu”)和种族侮辱性用语(如“chink”)相关的刻板印象普遍存在且具有语境敏感性。
- 发现语言模型会再现或反映其中许多刻板印象,证实了该数据集在评估模型行为方面的实用性。
- 尽管覆盖面广泛,由于抽样限制和刻板印象感知的主观性,该数据集仍不完整。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。