[论文解读] Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
本文研究了标注者身份与信念(尤其是政治保守主义和种族主义信念)如何在毒性语言检测中引入偏见。通过两项包含多样化参与者的在线研究,研究发现保守主义和种族偏见的标注者更少将针对黑人的语言评为有毒,但更可能将非洲裔美国人英语(AAE)标记为有毒,揭示了标注过程中系统性的主观性。研究进一步表明,像PerspectiveAPI这样的主流系统也反映了这些标注者的偏见,从而损害了自动化检测的公平性。
The perceived toxicity of language can vary based on someone's identity and beliefs, but this variation is often ignored when collecting toxic language datasets, resulting in dataset and model biases. We seek to understand the who, why, and what behind biases in toxicity annotations. In two online studies with demographically and politically diverse participants, we investigate the effect of annotator identities (who) and beliefs (why), drawing from social psychology research about hate speech, free speech, racist beliefs, political leaning, and more. We disentangle what is annotated as toxic by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity. Our results show strong associations between annotator identity and beliefs and their ratings of toxicity. Notably, more conservative annotators and those who scored highly on our scale for racist beliefs were less likely to rate anti-Black language as toxic, but more likely to rate AAE as toxic. We additionally present a case study illustrating how a popular toxicity detection system's ratings inherently reflect only specific beliefs and perspectives. Our findings call for contextualizing toxicity labels in social variables, which raises immense implications for toxic language annotation and detection.
研究动机与目标
- 调查标注者身份与信念如何影响语言标注中的毒性评分。
- 考察政治倾向、种族偏见和语言背景对文本冒犯性和种族主义感知的影响。
- 评估像PerspectiveAPI这样的主流毒性检测系统是否反映了特定标注者群体的偏见。
- 挑战NLP数据集与模型中毒性标签客观、普适的假设。
- 倡导在社会与人口变量的语境中对毒性标注进行重新审视。
提出的方法
- 通过两项在线研究,招募具有不同人口统计学和政治背景的参与者,收集对15个精选帖子和约600个帖子的毒性评分。
- 使用经过验证的心理量表测量标注者身份(种族、性别、政治倾向)和态度(种族主义信念、言论自由观点、传统主义)。
- 使用线性混合效应模型分析标注者特征与毒性评分之间的关系,同时控制个体差异。
- 计算PerspectiveAPI评分与标注者评分之间的皮尔逊相关系数,然后应用Fisher的r-to-z转换检验高分与低分标注者群体之间的显著差异。
- 分析三类文本:反黑人语言、非洲裔美国人英语(AAE)和粗俗语言,以区分偏见的来源。
- 使用效应量(Cohen's d,Pearson r)和经过Holm校正的p值来评估统计显著性与稳健性。
实验结果
研究问题
- RQ1标注者身份(如种族、性别、政治倾向)如何影响对文本毒性的评分?
- RQ2标注者信念——尤其是种族主义信念和言论自由观点——如何影响对冒犯性和种族主义的感知?
- RQ3不同文本特征(反黑人语言、AAE、粗俗语言)在多大程度上因标注者身份与信念而引发不同的毒性评分?
- RQ4现有的自动化毒性检测系统(如PerspectiveAPI)在多大程度上反映了特定标注者群体的偏见?
- RQ5聚合的毒性标签在哪些方面掩盖了毒性感知的主观性与社会语境?
主要发现
- 保守主义标注者显著更少将反黑人语言评为有毒,效应量在d = 0.34至d = 0.58之间(p < 0.001)。
- 在种族主义信念量表上得分较高的标注者更少将反黑人内容评为有毒(r = -0.21,p < 0.001),表明存在强烈负相关。
- 保守主义和传统主义标注者将AAE评为显著更具毒性的程度高于自由派或进步派标注者(d = 0.41,p < 0.001)。
- 种族主义信念水平较高的标注者更可能将AAE视为具有种族主义意图(r = 0.23,p < 0.001),表明将方言与恶意混为一谈。
- 保守主义和传统主义标注者将粗俗语言评为更具冒犯性(d = 0.38,p < 0.001),但这一差异在种族或性别上不显著。
- PerspectiveAPI的评分与种族主义信念水平较高的标注者在反黑人内容上的评分高度相关(r = 0.31,p < 0.001),表明该模型反映了这些偏见。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。