Skip to main content
QUICK REVIEW

[論文レビュー] Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?

Milagros Miceli, Julian Posada|arXiv (Cornell University)|Sep 16, 2021
Ethics and Social Impacts of AI参考文献 77被引用数 26
ひとこと要約

本論文は、偏りの緩和から、データ生産、労働、組織的文脈を検討するパワーアウェアなアプローチへの ML データ研究の転換を主張し、データ品質、データ労働、データ文書化実践の拡大を提案する。

ABSTRACT

Research in machine learning (ML) has primarily argued that models trained on incomplete or biased datasets can lead to discriminatory outputs. In this commentary, we propose moving the research focus beyond bias-oriented framings by adopting a power-aware perspective to "study up" ML datasets. This means accounting for historical inequities, labor conditions, and epistemological standpoints inscribed in data. We draw on HCI and CSCW work to support our argument, critically analyze previous research, and point at two co-existing lines of work within our community -- one bias-oriented, the other power-aware. This way, we highlight the need for dialogue and cooperation in three areas: data quality, data work, and data documentation. In the first area, we argue that reducing societal problems to "bias" misses the context-based nature of data. In the second one, we highlight the corporate forces and market imperatives involved in the labor of data workers that subsequently shape ML datasets. Finally, we propose expanding current transparency-oriented efforts in dataset documentation to reflect the social contexts of data design and production.

研究の動機と目的

  • 偏りに焦点を当てた枠組みは、MLデータ生産における権力ダイナミクスを見逃してしまう。
  • データ品質、データ労働、データ文書化を研究するためのパワーアウェアな視点を提唱する。
  • 労働条件と組織構造がデータセットと結果を形成する方法を強調する。
  • MLデータを“研究の上流”へと広げるための学際的対話を求める。

提案手法

  • バイアスに焦点を当てたMLデータ文献を批判的に分析し、HCI/CSCWのパワーアウェアな視点と対比させる。
  • データ労働の実践と文書化フレームワークの例を用いて、権力の非対称性がデータセットをいかに形成するかを示す。
  • MLデータを“研究の上流”へと持ち上げるための三つの方針(データ品質、データ労働、データ文書化)を提案する。
  • 研究を上流へと向かう概念(studying up、heteromation)を取り入れ、データの偏りをより広範な権力関係の症状として再解釈する。

実験結果

リサーチクエスチョン

  • RQ1組織内の権力の非対称性と労働実践は、MLデータの生産とデータセットにどのような影響を及ぼすか。
  • RQ2データセットの文書化をどのように拡張して、単なる偏りの緩和を超えた生産の文脈と権力動態を明らかにできるか。
  • RQ3データ労働者の条件とプラットフォーム統治はデータ品質と結果としてのMLシステムにどのように影響するか。
  • RQ4パワーアウェアなMLデータ研究を進展させる学際的な方法と協力関係は何か。

主な発見

  • 偏りの枠組みは、データセットに組み込まれた権力ダイナミクスと政治的作業を覆い隠す。
  • データ労働者の労働条件と組織構造は、データ品質とデータセットの結果に意味ある影響を与える。
  • 文書化フレームワークは、データセット構成だけでなく、生産文脈と権力関係を含むよう拡張できる。
  • パワーアウェアな分析は、力を持つ主体が支配する場合、デバイスが偏りを取り除いても不公正な結果を生む理由を明らかにする。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。