[论文解读] Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation
本研究探讨了评分者身份如何影响在线评论中的毒性标注,发现由自我认同为非裔美国人或LGBTQ的个体组成的专门评分者群体,在毒性评分上与对照组存在显著差异。基于身份特定评分者标注数据训练的模型,即使数据集较小,也优于基于随机评分者训练的模型,表明社区认同的标注者能产生更细致、更具包容性且偏差更小的模型结果。
Machine learning models are commonly used to detect toxicity in online conversations. These models are trained on datasets annotated by human raters. We explore how raters' self-described identities impact how they annotate toxicity in online comments. We first define the concept of specialized rater pools: rater pools formed based on raters' self-described identities, rather than at random. We formed three such rater pools for this study--specialized rater pools of raters from the U.S. who identify as African American, LGBTQ, and those who identify as neither. Each of these rater pools annotated the same set of comments, which contains many references to these identity groups. We found that rater identity is a statistically significant factor in how raters will annotate toxicity for identity-related annotations. Using preliminary content analysis, we examined the comments with the most disagreement between rater pools and found nuanced differences in the toxicity annotations. Next, we trained models on the annotations from each of the different rater pools, and compared the scores of these models on comments from several test sets. Finally, we discuss how using raters that self-identify with the subjects of comments can create more inclusive machine learning models, and provide more nuanced ratings than those by random raters.
研究动机与目标
- 检验评分者自我认同身份是否会影响针对种族与性取向相关评论的毒性标注。
- 评估基于身份群体(如非裔美国人和LGBTQ)的专门评分者群体是否能产生比随机评分者更细致、更少偏见的标注。
- 评估基于身份特定评分者标注数据训练的模型是否能优于基于随机标注者数据训练的模型。
- 创建并公开发布一个包含382,500条标注的全新数据集,来自三个经筛选的评分者群体,以支持未来研究。
- 倡导在机器学习流程中整合基于身份的标注者群体,以减少毒性检测中的偏见。
提出的方法
- 组建了三个专门的评分者群体:自我认同为非裔美国人的群体、LGBTQ群体以及两者皆非的群体——每个群体对同一组25,500条来自Civil Comments数据集的评论进行标注。
- 通过受控的身份导向评分者群体框架收集标注,确保自我认同的种族/性取向身份作为主要分组依据。
- 使用统计分析比较不同评分者群体之间的毒性评分,识别出评分差异与评分者身份之间的显著关联。
- 对评分差异较大的评论进行内容分析,识别导致分歧的语言和语境特征。
- 在每个评分者群体的标注数据上训练多种机器学习模型,并在多个测试集上评估其性能,以比较模型表现。
- 本研究发布了来自三个专门评分者群体的382,500条标注的新数据集,以支持可复现性研究和进一步研究。
实验结果
研究问题
- RQ1评分者的自我认同身份是否显著影响其对涉及种族与性取向评论的毒性标注?
- RQ2基于身份的评分者群体(非裔美国人、LGBTQ)的毒性评分与对照组(既非非裔美国人也非LGBTQ)的评分有何不同?
- RQ3即使数据集较小,基于专门评分者群体标注数据训练的模型是否仍能优于基于随机评分者数据训练的模型?
- RQ4哪些语言或语境特征可解释专门评分者群体与对照组之间评分差异的原因?
- RQ5如何通过基于身份的评分者群体提升毒性检测模型的公平性与包容性?
主要发现
- 评分者身份是毒性标注中的统计显著因素,非裔美国人和LGBTQ评分者对身份相关评论的评分与对照组存在显著差异。
- 在‘身份攻击’、‘威胁’和‘粗俗用语’类别中,专门评分者群体与对照组之间的评分差异最为显著。
- 基于专门评分者群体标注数据训练的模型,其性能优于基于更大规模随机评分者数据训练的模型,表明标注质量更高、更细致。
- 本研究公开发布了包含382,500条标注的全新数据集,来自三个专门评分者群体,以支持未来研究与模型评估。
- 初步内容分析显示,细微且具有社群特异性的语言标记(如微侵略)更准确地被身份特定评分者识别。
- 研究结果表明,引入社区认同的标注者可使毒性检测的机器学习模型更具包容性且偏差更小。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。