Skip to main content
QUICK REVIEW

[Paper Review] Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?

Milagros Miceli, Julian Posada|arXiv (Cornell University)|Sep 16, 2021
Ethics and Social Impacts of AISocial Sciences77 references26 citations
TL;DR

The paper argues shifting ML data research from bias mitigation to a power-aware approach that examines data production, labor, and organizational contexts, and proposes expanding data quality, data work, and data documentation practices.

ABSTRACT

Research in machine learning (ML) has primarily argued that models trained on incomplete or biased datasets can lead to discriminatory outputs. In this commentary, we propose moving the research focus beyond bias-oriented framings by adopting a power-aware perspective to "study up" ML datasets. This means accounting for historical inequities, labor conditions, and epistemological standpoints inscribed in data. We draw on HCI and CSCW work to support our argument, critically analyze previous research, and point at two co-existing lines of work within our community -- one bias-oriented, the other power-aware. This way, we highlight the need for dialogue and cooperation in three areas: data quality, data work, and data documentation. In the first area, we argue that reducing societal problems to "bias" misses the context-based nature of data. In the second one, we highlight the corporate forces and market imperatives involved in the labor of data workers that subsequently shape ML datasets. Finally, we propose expanding current transparency-oriented efforts in dataset documentation to reflect the social contexts of data design and production.

Motivation & Objective

  • Argue that bias-focused framing misses power dynamics in ML data production.
  • Advocate for a power-aware lens to study data quality, data work, and data documentation.
  • Highlight how labor conditions and organizational structures shape datasets and outcomes.
  • Call for interdisciplinary dialogue between CS, sociology, anthropology, and economics to study up ML data.

Proposed method

  • Critically analyze bias-focused ML data literature and contrast it with power-aware perspectives from HCI/CSCW.
  • Use examples from data work practices and documentation frameworks to illustrate how power asymmetries shape datasets.
  • Propose a three-pronged agenda (data quality, data work, data documentation) to study up ML data.
  • Draw on interdisciplinary concepts (studying up, heteromation) to reframe data bias as a symptom of broader power relations.

Experimental results

Research questions

  • RQ1How do power asymmetries within organizations and labor practices influence ML data production and datasets?
  • RQ2In what ways can dataset documentation be expanded to reveal production contexts and power dynamics beyond mere bias mitigation?
  • RQ3How do data workers’ conditions and platform governance affect data quality and the resulting ML systems?
  • RQ4What interdisciplinary methods and collaborations can advance a power-aware study of ML data?

Key findings

  • Bias framing obscures power dynamics and political work embedded in datasets.
  • Data workers’ labor conditions and organizational structures meaningfully shape data quality and dataset outcomes.
  • Documentation frameworks can be extended to include production contexts and power relations, not just dataset composition.
  • Power-aware analyses can reveal why debiased data may still produce unjust outcomes when controlled by powerful actors.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.