[论文解读] How True is GPT-2? An Empirical Analysis of Intersectional Occupational Biases.
本文通过将性别与宗教、性取向、种族、政治归属及姓名来源交叉分析,实证研究了 GPT-2 的职业偏见,基于 396,000 个句子补全结果。研究发现,GPT-2 对女性生成的职业关联更刻板且多样性更低,尤其在交叉身份情境下更为明显,并揭示了受保护类别之间存在显著的交互效应,引发了关于语言模型应学习何种职业关联的规范性问题。
The capabilities of natural language models trained on large-scale data have increased immensely over the past few years. Downstream applications are at risk of inheriting biases contained in these models, with potential negative consequences especially for marginalized groups. In this paper, we analyze the occupational biases of a popular generative language model, GPT-2, intersecting gender with five protected categories: religion, sexuality, ethnicity, political affiliation, and name origin. Using a novel data collection pipeline we collect 396k sentence completions of GPT-2 and find: (i) The machine-predicted jobs are less diverse and more stereotypical for women than for men, especially for intersections; (ii) Fitting 262 logistic models shows intersectional interactions to be highly relevant for occupational associations; (iii) For a given job, GPT-2 reflects the societal skew of gender and ethnicity in the US, and in some cases, pulls the distribution towards gender parity, raising the normative question of what language models _should_ learn.
研究动机与目标
- 探究 GPT-2 如何将职业与交叉身份(包括性别、宗教、性取向、种族、政治归属及姓名来源)关联起来。
- 评估该模型是否放大或缓解了社会中的职业刻板印象,尤其是对边缘化群体的影响。
- 评估交叉身份交互作用在多大程度上塑造了 GPT-2 的职业预测。
提出的方法
- 开发了一种新型数据收集管道,利用嵌入交叉身份的提示词,从 GPT-2 生成了 396,000 个句子补全结果。
- 通过在提示中嵌入身份信息(如 '她是一位 [职业]'),其中主体的身份由多个属性共同定义。
- 对 262 种不同的职业类别拟合逻辑回归模型,以量化受保护属性的个体效应与交互效应。
- 将 GPT-2 的预测职业分布与真实世界中的美国职业人口统计数据进行比较,以评估其一致性与偏差程度。
- 通过检验逻辑模型中交互项的显著性与方向,衡量交叉偏见。
实验结果
研究问题
- RQ1当性别与其他受保护类别交叉时,GPT-2 对男性与女性的职业预测有何差异?
- RQ2交叉身份交互作用(如女性 + 穆斯林 + 黑人)在多大程度上显著影响 GPT-2 的职业预测?
- RQ3GPT-2 的职业分布与美国真实世界中的性别与种族职业构成在多大程度上一致?
- RQ4GPT-2 的输出是否在某些职业中趋向性别平等?如果是,其条件是什么?
- RQ5语言模型学习到扭曲或追求平等的职业关联,会引发哪些规范性影响?
主要发现
- 与男性相比,GPT-2 对女性生成的职业关联显著更少多样化且更具性别刻板印象,尤其在交叉身份情境下更为明显。
- 性别与宗教或种族等属性的交叉交互作用在塑造职业预测方面具有高度显著性,表明存在非加法性的偏见效应。
- 对于许多职业,GPT-2 的预测性别分布与美国真实职业人口统计数据一致,反映出社会中的结构性偏差。
- 在某些情况下,GPT-2 的预测使分布趋向性别平等,表明模型行为并不总是反映现实中的不平等。
- 该模型的行为引发了规范性担忧:语言模型是否应反映现有不平等,还是应致力于实现职业关联中的公平性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。