[Paper Review] Data and its (dis)contents: A survey of dataset development and use in machine learning research
This paper critically examines the role of datasets in machine learning research, arguing that current practices in data collection, annotation, and benchmarking perpetuate biases, spurious correlations, and ethical issues. It advocates for a paradigm shift toward more careful, context-aware, and ethically responsible dataset development that prioritizes representativeness, transparency, and respect for data subjects over scale and performance metrics.
Datasets have played a foundational role in the advancement of machine learning research. They form the basis for the models we design and deploy, as well as our primary medium for benchmarking and evaluation. Furthermore, the ways in which we collect, construct and share these datasets inform the kinds of problems the field pursues and the methods explored in algorithm development. However, recent work from a breadth of perspectives has revealed the limitations of predominant practices in dataset collection and use. In this paper, we survey the many concerns raised about the way we collect and use data in machine learning and advocate that a more cautious and thorough understanding of data is necessary to address several of the practical and ethical issues of the field.
Motivation & Objective
- To identify and analyze systemic flaws in dataset design and use that undermine the validity and ethics of machine learning research.
- To highlight how current data collection practices—especially web scraping and crowdwork—mask human labor, bias, and contextual dependencies.
- To critique the overreliance on benchmark datasets as drivers of research progress, which often prioritize performance over real-world relevance and fairness.
- To advocate for a cultural shift in ML research toward datasets that are contextually grounded, ethically sourced, and transparently documented.
- To emphasize the need for broader evaluation frameworks beyond benchmarking to support equitable and responsible AI development.
Proposed method
- Conducting a comprehensive survey of recent literature on dataset-related issues in NLP and computer vision.
- Categorizing critiques into four themes: representational bias, spurious correlations, flawed task framing, and poor documentation and annotation practices.
- Analyzing case studies of problematic datasets (e.g., ImageNet, OntoNotes, toxicity datasets) to illustrate systemic issues in data construction.
- Evaluating proposed technical solutions such as adversarial datasets and data augmentation, while critiquing their limitations in addressing root causes.
- Surveying broader institutional and cultural critiques of data reuse, legal risks, and data management practices in ML research.
- Advocating for a research culture that values context, consent, and interdisciplinary collaboration over scale and leaderboard performance.
Experimental results
Research questions
- RQ1How do representational biases in machine learning datasets reflect and reinforce societal inequities?
- RQ2To what extent do spurious correlations in benchmark datasets enable models to 'game' tasks without learning meaningful capabilities?
- RQ3Why is the current benchmark-driven culture in machine learning research problematic for scientific progress and ethical deployment?
- RQ4What are the ethical and legal risks associated with large-scale web scraping and reuse of data without consent?
- RQ5How can dataset development be reformed to prioritize context, transparency, and respect for data subjects?
Key findings
- Prominent datasets like ImageNet and OntoNotes exhibit significant underrepresentation of marginalized sociodemographic groups, including darker-skinned individuals and female pronouns.
- Datasets frequently encode harmful stereotypes—such as gendered associations between occupations and gender in vision and language data—leading to biased model behavior.
- The ImageNet dataset was found to contain millions of images labeled with racial slurs and derogatory terms, prompting partial removal of the dataset.
- Many benchmark datasets are gameable due to spurious correlations (e.g., text containing 'gay' being labeled as toxic), undermining claims of model generalization.
- Current data collection practices often obscure the labor, context, and subjectivity involved in dataset creation, leading to a lack of transparency and accountability.
- Efforts to fix datasets post-hoc—such as adversarial data creation or filtering—fail to resolve deeper issues of representation, context, and ethical sourcing.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.