[Paper Review] Bringing the People Back In: Contesting Benchmark Machine Learning Datasets
The paper proposes a genealogical research program to study benchmark ML datasets as infrastructural artifacts, aiming to reveal histories, values, and labor behind dataset construction and to enable contestability rather than mere transparency.
In response to algorithmic unfairness embedded in sociotechnical systems, significant attention has been focused on the contents of machine learning datasets which have revealed biases towards white, cisgender, male, and Western data subjects. In contrast, comparatively less attention has been paid to the histories, values, and norms embedded in such datasets. In this work, we outline a research program - a genealogy of machine learning data - for investigating how and why these datasets have been created, what and whose values influence the choices of data to collect, the contextual and contingent conditions of their creation. We describe the ways in which benchmark datasets in machine learning operate as infrastructure and pose four research questions for these datasets. This interrogation forces us to "bring the people back in" by aiding us in understanding the labor embedded in dataset construction, and thereby presenting new avenues of contestation for other researchers encountering the data.
Motivation & Objective
- Motivate a genealogical method to study how benchmark ML datasets are created and what values influence data collection.
- Frame datasets as infrastructure that shapes research agendas, benchmarks, and industry practice.
- Introduce a vocabulary and analytical lens from infrastructural studies to de-naturalize data practices.
- Outline a four-part research program to understand motivations, histories, authority, and current practices surrounding benchmark datasets.
Proposed method
- Adopt Michel Foucault’s genealogy to trace historical formation and transformation of dataset practices.
- Use infrastructural inversion to reveal the hidden labor and contextual factors in data creation.
- Treat datasets and benchmarks as infrastructure that undergirds ML research and industry deployment.
- Apply text analysis of dataset documentation and related communications to uncover motivations and conventions.
- Propose ethnographic, historical, and multi-site inquiry to study data work practices in major ML hubs.
Experimental results
Research questions
- RQ1How do dataset developers describe and motivate the decisions involved in creating datasets and their documentation?
- RQ2What are the histories and contingent conditions of creation of benchmark datasets in machine learning?
- RQ3How do benchmark datasets become authoritative, and how does this authority shape research practice and norms?
- RQ4What are the current work practices, norms, and routines that structure data collection, curation, and annotation in machine learning?
Key findings
- Introduces new vocabulary and concepts from infrastructural studies to frame data as power-laden infrastructure and to encourage contestability.
- Outlines a novel genealogy of machine learning data as a research program with explicit questions and methods.
- Argues that controlling the data pipeline requires examining historical contingencies, power relations, and labor involved in dataset creation.
- Advocates for data release practices that document objectives, collection methodologies, curation, and classification to support reflexive analysis.
- Emphasizes moving beyond data quantity as the sole solution to fairness, highlighting the risks of predatory inclusion and data labor exploitation.
- Proposes in situ, multi-sited ethnography of major ML hubs to uncover current data practices and normative routines.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.