Skip to main content
QUICK REVIEW

[论文解读] Extending Challenge Sets to Uncover Gender Bias in Machine Translation: Impact of Stereotypical Verbs and Adjectives

Jonas Troles, Ute Schmid|ArXiv.org|Jul 24, 2021
Hate Speech and Cyberbullying Detection参考文献 13被引用 6
一句话总结

本文将 WinoMT 挑战集扩展为 WiBeMT,构建了一个包含超过 70,000 个英德翻译句子的大规模评估集,其中包含性别刻板印象的形容词和动词,以揭示神经机器翻译(NMT)中的性别偏见。研究发现,形容词显著影响翻译中的性别偏见,所有三种商用机器翻译系统(DeepL、Microsoft、Google 翻译)均表现出偏差,尽管 DeepL 的偏差最低。研究结果表明,性别化的形容词在某些情况下反而可能减少性别歧视,但也可能引入新的偏见形式。

ABSTRACT

Human gender bias is reflected in language and text production. Because state-of-the-art machine translation (MT) systems are trained on large corpora of text, mostly generated by humans, gender bias can also be found in MT. For instance when occupations are translated from a language like English, which mostly uses gender neutral words, to a language like German, which mostly uses a feminine and a masculine version for an occupation, a decision must be made by the MT System. Recent research showed that MT systems are biased towards stereotypical translation of occupations. In 2019 the first, and so far only, challenge set, explicitly designed to measure the extent of gender bias in MT systems has been published. In this set measurement of gender bias is solely based on the translation of occupations. In this paper we present an extension of this challenge set, called WiBeMT, with gender-biased adjectives and adds sentences with gender-biased verbs. The resulting challenge set consists of over 70, 000 sentences and has been translated with three commercial MT systems: DeepL Translator, Microsoft Translator, and Google Translate. Results show a gender bias for all three MT systems. This gender bias is to a great extent significantly influenced by adjectives and to a lesser extent by verbs.

研究动机与目标

  • 为解决现有机器翻译性别偏见评估范围有限的问题,现有研究仅关注职业术语。
  • 探究性别刻板印象的形容词和动词如何影响神经机器翻译系统中的翻译结果。
  • 开发一个更大、更具多样性的挑战集(WiBeMT),以捕捉超越职业术语的更广泛的语言性别偏见来源。
  • 使用扩展的数据集评估三种商用 NMT 系统(DeepL、Microsoft 翻译和 Google 翻译)的性别偏见。
  • 评估性别化的形容词和动词是否在翻译输出中加剧或缓解性别偏见。

提出的方法

  • 通过使用词嵌入和与性别特定词表的余弦相似度,识别出具有性别刻板印象的形容词和动词,从而扩展 WinoBias 数据集。
  • 通过将职业术语与刻板印象形容词和动词结合,构建新句子,形成性别一致和不一致的配对。
  • 共生成 70,686 个句子,系统性地变换代词、形容词和动词,以测试翻译中的性别偏见。
  • 使用三种商用 NMT 系统(DeepL、Microsoft 翻译和 Google 翻译)对所有句子进行翻译。
  • 使用 %TCG(性别代词翻译正确率)等指标评估性别偏见,并比较男性、女性和中性条件下的准确率差异。
  • 分析形容词和动词对翻译性别结果的影响,区分其对职业名词翻译和整体系统偏见的影响。

实验结果

研究问题

  • RQ1性别刻板印象的形容词在德语-英语翻译中如何影响机器翻译输出的性别偏见?
  • RQ2性别化的动词在多大程度上导致神经机器翻译系统中的偏差翻译?
  • RQ3在包含性别化形容词的句子中,性别代词引用的翻译准确率差异是否减少或增加?
  • RQ4在处理具有刻板印象形容词和动词的句子时,商用 NMT 系统(DeepL、Microsoft、Google 翻译)在性别偏见方面如何比较?
  • RQ5尽管引入了新的偏见,添加性别化形容词是否可能反而减少整体性别偏见?

主要发现

  • 形容词显著影响 NMT 系统中的性别偏见,无论是女性还是男性形容词都会使翻译偏向女性性别,尽管女性形容词的影响更强。
  • DeepL 翻译在三种系统中性别偏见最低,其在男性和女性代词句子中的 %TCG 差异最小。
  • Google 翻译整体表现出对男性翻译的最高偏好,尤其在动词句中,其表现甚至低于 Microsoft 翻译。
  • 尽管整体存在偏见,但添加性别化形容词可提高女性代词句子的翻译准确率,从而缩小性别准确率差距。
  • 性别化动词对翻译偏见的影响小于形容词,且仅在 DeepL 和 Microsoft 翻译中具有显著影响,Google 翻译中则不显著。
  • 研究揭示了一种悖论性效应:尽管形容词增加了女性性别翻译的可能性,但它们通过提高女性指代的翻译准确率,反而减少了系统输出的整体性别歧视。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。