[论文解读] Semantics derived automatically from language corpora necessarily contain human biases.
本文表明,像 GloVe 这类在大规模网络文本语料上训练的词嵌入模型,会自动学习并复制人类类似的语义偏见——例如与种族、性别和社会角色相关的偏见——因为这些偏见已内嵌于语言本身。通过使用新颖的评估工具(WEAT 和 WEFAT),研究发现统计机器学习模型继承社会偏见并非出于设计,而是通过接触有偏见的语言所致,揭示了语言语料库中编码了历史和文化偏见,而这些偏见随后被人工智能系统所捕获。
Artificial intelligence and machine learning are in a period of astounding growth. However, there are concerns that these technologies may be used, either with or without intention, to perpetuate the prejudice and unfairness that unfortunately characterizes many human institutions. Here we show for the first time that human-like semantic biases result from the application of standard machine learning to ordinary language---the same sort of language humans are exposed to every day. We replicate a spectrum of standard human biases as exposed by the Implicit Association Test and other well-known psychological studies. We replicate these using a widely used, purely statistical machine-learning model---namely, the GloVe word embedding---trained on a corpus of text from the Web. Our results indicate that language itself contains recoverable and accurate imprints of our historic biases, whether these are morally neutral as towards insects or flowers, problematic as towards race or gender, or even simply veridical, reflecting the status quo for the distribution of gender with respect to careers or first names. These regularities are captured by machine learning along with the rest of semantics. In addition to our empirical findings concerning language, we also contribute new methods for evaluating bias in text, the Word Embedding Association Test (WEAT) and the Word Embedding Factual Association Test (WEFAT). Our results have implications not only for AI and machine learning, but also for the fields of psychology, sociology, and human ethics, since they raise the possibility that mere exposure to everyday language can account for the biases we replicate here.
研究动机与目标
- 调查在日常语言上训练的机器学习模型是否会继承类似人类的语义偏见。
- 检验广泛使用的自然语言处理模型(如 GloVe)是否反映心理学研究(如内隐联想测试)中已知的心理偏见。
- 开发用于检测词嵌入中偏见的新评估方法。
- 证明语言语料中的偏见足以在人工智能模型中产生有偏见的语义表征。
提出的方法
- 在大规模网络文本语料上训练 GloVe 词嵌入模型,以学习词语的稠密向量表征。
- 应用词嵌入关联测试(WEAT)来测量词类(如种族、性别)与属性(如愉悦/不愉悦)之间的关联。
- 使用词嵌入事实关联测试(WEFAT)来评估词语与事实性社会分布之间的关联(如性别与职业、名字与性别)。
- 通过词嵌入中的统计关联复制内隐联想测试中已知的心理学偏见。
- 将模型推导出的关联与人类观察到的偏见进行比较,以验证语义偏见的复制效果。
- 分析不同词嵌入维度和语义类别中偏见模式的一致性与准确性。
实验结果
研究问题
- RQ1在网页文本上训练的词嵌入在多大程度上再现了心理学研究中已知的人类语义偏见?
- RQ2像 GloVe 这类标准自然语言处理模型是否能自动学习并反映语言语料中存在的社会偏见?
- RQ3词嵌入中性别与职业之间的关联与现实世界的人口统计数据分布相比如何?
- RQ4WEAT 和 WEFAT 框架能否可靠地检测并量化词嵌入中的偏见?
- RQ5仅通过接触自然语言(而无需明确指令)是否会导致机器学习模型内化社会偏见?
主要发现
- 在网页文本上训练的词嵌入复制了广泛的人类类似偏见,包括与种族、性别和社会角色相关的偏见,这些偏见通过 WEAT 测量得到证实。
- 该研究证实,诸如‘护士’与‘女性’、‘工程师’与‘男性’之间的性别关联等偏见,被 GloVe 模型准确捕捉。
- WEFAT 测试显示,词嵌入能以高精度反映事实性人口统计数据,如性别化姓名和职业性别比例。
- 偏见的复制并非源于模型设计,而是源自训练语言数据中存在的统计规律性。
- 结果表明,即使像‘花朵’与‘愉悦’这类道德中立的关联,也编码在嵌入中,表明偏见是基于语言的人工智能系统中的系统性特征。
- 研究结果表明,语言本身是将社会偏见嵌入机器学习系统的主要载体。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。