Skip to main content
QUICK REVIEW

[论文解读] Toxic Bias: Perspective API Misreads German as More Toxic

Gianluca Nogara, Francesco Pierri|arXiv (Cornell University)|Dec 19, 2023
Hate Speech and Cyberbullying Detection被引用 5
一句话总结

本文識別出 Google 的 Perspective API 存在系統性偏見,即德語內容在整體上被錯誤分類為更具毒性,其毒性分數平均高於英語翻譯版本的四倍。透過多語言的 Twitter 和 Wikipedia 數據集,作者證明此偏見在不同主題與資料來源間均具穩定性,因模型的不透明與專有性,引發對不公平內容審查與計算社會科學研究偏誤的擔憂。

ABSTRACT

Proprietary public APIs play a crucial and growing role as research tools among social scientists. Among such APIs, Google's machine learning-based Perspective API is extensively utilized for assessing the toxicity of social media messages, providing both an important resource for researchers and automatic content moderation. However, this paper exposes an important bias in Perspective API concerning German language text. Through an in-depth examination of several datasets, we uncover intrinsic language biases within the multilingual model of Perspective API. We find that the toxicity assessment of German content produces significantly higher toxicity levels than other languages. This finding is robust across various translations, topics, and data sources, and has significant consequences for both research and moderation strategies that rely on Perspective API. For instance, we show that, on average, four times more tweets and users would be moderated when using the German language compared to their English translation. Our findings point to broader risks associated with the widespread use of proprietary APIs within the computational social sciences.

研究动机与目标

  • 調查 Perspective API 是否在毒性分數上對德語表現出語言特定的偏見,特別是針對德語。
  • 評估此類偏見對多語言線上平台研究與內容審查的影響。
  • 強調在計算社會科學中依賴專有、黑箱人工智慧模型所帶來的風險。
  • 呼籲提升人工智慧驅動內容審查系統的透明度與責任制。
  • 證明相同內容在德語中獲得的毒性分數顯著高於英語版本,從而破壞公平與正義。

提出的方法

  • 作者分析了多語言的 Twitter 討論與隨機 Wikipedia 摘要,以比較不同語言之間的毒性分數。
  • 他們將相同內容翻譯為德語與英語,並測量 Perspective API 在毒性分數上的差異。
  • 統計分析用於評估德語內容毒性分數較高的顯著性,涵蓋多種資料集與主題。
  • 本研究依賴 Perspective API 的公開毒性分數系統,該系統針對多種毒性類別輸出 0 到 1 之間的數值。
  • 研究人員使用大規模資料集,以確保研究結果在多樣的語言與語境設定下具備穩健性與普遍性。
  • 由於模型的專有性,無法對模型架構或訓練資料進行內部分析。
Figure 1. Distribution of toxicity scores for tweets shared in German-speaking countries (Austria, Switzerland, and Germany) versus those in other EU countries. Distributions are statistically different at $\alpha=0.05$ according to a Kruskal-Wallis test. Median toxicity is 0.075 for German-speaking
Figure 1. Distribution of toxicity scores for tweets shared in German-speaking countries (Austria, Switzerland, and Germany) versus those in other EU countries. Distributions are statistically different at $\alpha=0.05$ according to a Kruskal-Wallis test. Median toxicity is 0.075 for German-speaking

实验结果

研究问题

  • RQ1Perspective API 是否對德語內容賦予顯著較高的毒性分數,相比其他語言?
  • RQ2這些較高的分數是否在不同主題、資料來源與翻譯方法間保持一致?
  • RQ3此偏見對依賴 Perspective API 進行跨語言毒性分析的學術研究有何影響?
  • RQ4此偏見如何影響多語言平台的內容審查實務?
  • RQ5API 的黑箱性質在多大程度上阻礙研究人員檢測或修正此類語言特定偏見?

主要发现

  • 德語內容獲得的毒性分數,平均為其英語翻譯版本的四倍。
  • 此偏見在多個資料集(包括多語言 Twitter 討論與 Wikipedia 摘要)中一致出現。
  • 德語內容的毒性分數呈現尖峰分佈,顯示模型人工製品或錯誤,而非自然語言差異。
  • 此偏見在不同主題、內容類型與資料來源下均持續存在,顯示此為多語言模型內在的系統性問題。
  • 研究結果顯示,德語使用者面臨被自動審查系統錯誤標記或審查的不成比例風險。
  • 本研究強調在缺乏透明度或驗證的情況下,依賴專有且不透明的人工智慧工具進行研究與平台治理所帶來的風險。
Figure 2. Top 10 most frequent Perspective API scores rounded at 8 decimal digits for tweets shared in German-speaking countries (Austria, Switzerland and Germany) and those in other EU countries, from Dataset 1.
Figure 2. Top 10 most frequent Perspective API scores rounded at 8 decimal digits for tweets shared in German-speaking countries (Austria, Switzerland and Germany) and those in other EU countries, from Dataset 1.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。