[論文レビュー] Assessing Gender Bias in Machine Translation -- A Case Study with Google Translate
本研究は、12ヶ国語から英語への翻訳を通じて、Google Translateにおける性別バイアスを調査した。性別に中立な文(例:「彼/彼女はエンジニアです」)を翻訳した結果、特にSTEM職において、実際の職業における性別分布とは対照的に、男性を前提とする傾向が強く見られた。著者らは、Google Translateが翻訳において女性を体系的に低く評価しており、人口統計データを超えたバイアスを示していることを示し、自然言語処理システムにおけるバイアス除去技術の導入を提唱している。
Recently there has been a growing concern about machine bias, where trained statistical models grow to reflect controversial societal asymmetries, such as gender or racial bias. A significant number of AI tools have recently been suggested to be harmfully biased towards some minority, with reports of racist criminal behavior predictors, Iphone X failing to differentiate between two Asian people and Google photos' mistakenly classifying black people as gorillas. Although a systematic study of such biases can be difficult, we believe that automated translation tools can be exploited through gender neutral languages to yield a window into the phenomenon of gender bias in AI. In this paper, we start with a comprehensive list of job positions from the U.S. Bureau of Labor Statistics (BLS) and used it to build sentences in constructions like "He/She is an Engineer" in 12 different gender neutral languages such as Hungarian, Chinese, Yoruba, and several others. We translate these sentences into English using the Google Translate API, and collect statistics about the frequency of female, male and gender-neutral pronouns in the translated output. We show that GT exhibits a strong tendency towards male defaults, in particular for fields linked to unbalanced gender distribution such as STEM jobs. We ran these statistics against BLS' data for the frequency of female participation in each job position, showing that GT fails to reproduce a real-world distribution of female workers. We provide experimental evidence that even if one does not expect in principle a 50:50 pronominal gender distribution, GT yields male defaults much more frequently than what would be expected from demographic data alone. We are hopeful that this work will ignite a debate about the need to augment current statistical translation tools with debiasing techniques which can already be found in the scientific literature.
研究の動機と目的
- 機械翻訳システム(Google Translateなど)が、人称代名詞の選択において性別バイアスを示すかどうかを調査すること。
- このようなバイアスが、職業における実際の性別分布を反映しているのか、それ以上に顕著であるのかを検討すること。
- 言語のタイプ(例:性別を含む言語 vs. 性別を含まない言語)が翻訳バイアスに与える影響を評価すること。
- 自動翻訳システムを通じて、社会的な性別ステレオタイプを検出する可能性を評価すること。
- 統計的機械翻訳パイプラインにバイアス除去技術を統合するよう提唱すること。
提案手法
- 米国労働統計局(BLS)の職業名を用いて、性別に中立な文を構築し、『彼/彼女は[職業]です』の形式とした。
- Google Translate APIを用い、12ヶ国語(例:ハンガリー語、中国語、ヨルーバ語)の性別に中立な言語から英語に翻訳した。
- 翻訳出力における男性代名詞、女性代名詞、性別に中立な代名詞の頻度を収集・分析した。
- 各職業における実際の女性の参加率(BLSデータ)と照らし合わせ、翻訳における女性代名詞の頻度を比較した。
- 『控えめな』や『罪深い』といった形容詞を用い、性別に偏った形容詞が翻訳における人称代名詞の選択に与える影響をテストした。
- ハンガリー語、中国語、バスク語などの言語間での翻訳出力を比較することで、言語固有のバイアスの違いを評価した。
実験結果
リサーチクエスチョン
- RQ1Google Translateは、性別に中立な文を英語に翻訳する際、人称代名詞の選択において男性を前提とする傾向を示すか?
- RQ2翻訳における女性代名詞の頻度は、各職業における実際の女性参加率とどの程度相関しているか?
- RQ3特に性別に中立な言語を含む、さまざまな出力言語(特に性別に中立な言語)が、翻訳出力における性別バイアスの度合いにどのように影響するか?
- RQ4性別に偏った形容詞(例:『控えめな』、『罪深い』)は、翻訳において女性または男性の代名詞を不釣り合いに多く使用させるか?
- RQ5機械翻訳の出力は、学習データに埋め込まれた社会的な性別ステレオタイプを検出するための代理指標として機能できるか?
主な発見
- Google Translateは、特にSTEM職の職業名において、性別に中立な言語からでも強い男性を前提とする傾向を示している。
- 翻訳における男性代名詞の頻度は、これらの職業における実際の女性参加率に基づく予測をはるかに上回っている。
- 例えば、女性の参加率が低い職業(例:『ソフトウェア開発者』)においても、Google Translateは予想を上回る割合の男性代名詞を出力した。
- 『控えめな』や『魅力的な』といった形容詞は、女性代名詞の割合が高くなる傾向を示したが、一方で『罪深い』や『冷酷な』といった形容詞は、ほとんどが男性代名詞と結びついていた。
- ハンガリー語では中国語よりも代名詞の分布がよりバランスが取れており、言語構造がバイアスの伝播に影響を与えることが示された。
- バスク語やヨルーバ語はしばしば性別に中立な翻訳を出力したが、バスク語では識別不能な性別代名詞の割合も高く、言語固有のバイアス検出の課題が浮き彫りになった。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。