[論文レビュー] Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models
本論文は、性別と宗教、性的指向、民族、政治的信条、そして大陸名の起源の交差を横断する職業連想の出荷時設定GPT-2バイアスを分析し、予測を米国の労働市場データと比較します。
The capabilities of natural language models trained on large-scale data have increased immensely over the past few years. Open source libraries such as HuggingFace have made these models easily available and accessible. While prior research has identified biases in large language models, this paper considers biases contained in the most popular versions of these models when applied `out-of-the-box' for downstream tasks. We focus on generative language models as they are well-suited for extracting biases inherited from training data. Specifically, we conduct an in-depth analysis of GPT-2, which is the most downloaded text generation model on HuggingFace, with over half a million downloads per month. We assess biases related to occupational associations for different protected categories by intersecting gender with religion, sexuality, ethnicity, political affiliation, and continental name origin. Using a template-based data collection pipeline, we collect 396K sentence completions made by GPT-2 and find: (i) The machine-predicted jobs are less diverse and more stereotypical for women than for men, especially for intersections; (ii) Intersectional interactions are highly relevant for occupational associations, which we quantify by fitting 262 logistic models; (iii) For most occupations, GPT-2 reflects the skewed gender and ethnicity distribution found in US Labor Bureau data, and even pulls the societally-skewed distribution towards gender parity in cases where its predictions deviate from real labor market observations. This raises the normative question of what language models should learn - whether they should reflect or correct for existing inequalities.
研究の動機と目的
- 出荷時点で使用される人気のオープンソース言語モデルにおける偏見の研究を動機づける。
- 複数の保護属性にわたるGPT-2の交差的職業バイアスを測定するためのプロトコルを開発する。
- 統計的モデリングを通じて、交差性が職業予測に与える影響を定量化する。
- GPT-2 が予測する職業分布を実データの米国労働市場データと比較し、整合性または乖離を評価する。
提案手法
- 職業名を含むGPT-2文の396K件を抽出する、テンプレートベースのデータ収集パイプラインを用いる。
- アイデンティティベースおよび名前ベースの接頭辞プロンプトを適用して、保護属性に関連する職業連想を引き出す。
- Stanford CoreNLP NERを用いて職業を抽出し、分析のためのワンホットトークン頻度マトリクスを作成する。
- 交互作用項を含む262個のロジスティック回帰モデルを適合させ、職業予測における交差効果を定量化する。
- GPT-2予測分布を2019年の米国労働統計局の職業データと比較し、人口統計分布を調整する。
実験結果
リサーチクエスチョン
- RQ1出荷時設定のGPT-2出力は性別および交差的職業バイアスを示すか。
- RQ2性別と民族、宗教、性的指向、政治的信条、大陸名起源との交差は、予測される職業にどのように影響するか。
- RQ3性別と人種/民族別に、GPT-2の予測が現実の米国の職業分布とどの程度一致しているか。
- RQ4主効果を超えたGPT-2の職業予測の変化に、交差的相互作用は有意か。
- RQ5モデルの予測は職業連想における表象的または割当的な害を示すか。
主な発見
- GPT-2は女性に対して男性よりも多様性が低く、より固定観念的な職の集合を返し、女性の職業集積が高い。
- 交差的相互作用(例:性別と民族、宗教、性的指向)は職業予測を大きく形作り、多くのモデルで有意な相互作用項を示す。
- ほとんどの職業でGPT-2の予測は米国労働データの歪みを反映しており、場合によっては男女平等へ傾くことで、極端な現実データの歪みを是正する可能性がある。
- ロジスティック回帰では女性ダミーが多くのモデルで説明可能な変動を加え(平均ΔR^2 約+3.3%)、多くの相互作用が有意である(回帰の約3分の1)。
- 米国データとの比較は強い正の関連を示し(ベースライン性別で Kendall Tau ≈ 0.63)、平均二乗誤差は小さいが、GPT-2は女性が多い職を過大評価し、女性比率が高い職業では過小評価することが多い。
- GPT-2は女性の上位5職(例:ウェイトレス、看護師)を米国データに比べ過剰表示する傾向がある一方、民族性に関しては現実の分布を一部反映している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。