[Paper Review] Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators
This paper presents insights from 18 dataset creators and articulates seven practical recommendations to improve responsible dataset design, focusing on data quality, diversity, documentation, and governance.
The increasing demand for high-quality datasets in machine learning has raised concerns about the ethical and responsible creation of these datasets. Dataset creators play a crucial role in developing responsible practices, yet their perspectives and expertise have not yet been highlighted in the current literature. In this paper, we bridge this gap by presenting insights from a qualitative study that included interviewing 18 leading dataset creators about the current state of the field. We shed light on the challenges and considerations faced by dataset creators, and our findings underscore the potential for deeper collaboration, knowledge sharing, and collective development. Through a close analysis of their perspectives, we share seven central recommendations for improving responsible dataset creation, including issues such as data quality, documentation, privacy and consent, and how to mitigate potential harms from unintended use cases. By fostering critical reflection and sharing the experiences of dataset creators, we aim to promote responsible dataset creation practices and develop a nuanced understanding of this crucial but often undervalued aspect of machine learning research.
Motivation & Objective
- Highlight the role and perspectives of dataset creators in responsible ML practices.
- Identify common challenges across diverse datasets and institutional contexts.
- Propose actionable recommendations to improve data quality, diversity, consent, and use limitations.
- Encourage collaboration and professionalization of data work within the ML community.
Proposed method
- Conducted open-ended qualitative interviews with 18 dataset creators (July–Sept 2022).
- Recruited 47 potential participants; 18 participated (38% response rate).
- Used semi-structured interviews to explore dataset origins, usage, maintenance, and obsolescence.
- Transcribed interviews and performed iterative thematic coding to extract recommendations.
- Ensured anonymity or identification choices for participants; study approved by IRB.
Experimental results
Research questions
- RQ1What are the challenges and best practices faced by dataset creators across domains?
- RQ2What concrete recommendations do creators offer to improve responsible dataset creation?
- RQ3How can the dataset community move toward better collaboration and professionalization?
- RQ4What role do documentation, diversity, validation, and consent play in responsible dataset design?
- RQ5How should datasets be communicated, used, and potentially retired or updated over time?
Key findings
- Seven central recommendations for responsible dataset creation emerged from interviews.
- Diversity and thorough auditing are crucial to mitigate bias and unintended harms.
- High data quality requires careful validation, manual inspection, and curation, with awareness of trade-offs.
- Early, iterative development with learning from mistakes is essential.
- Open documentation and clear communication of limitations support reproducibility and reuse.
- Datasets should be user-centric with clearly defined intended uses and consideration of unintended use cases.
- Ethical considerations, consent, privacy, licensing, attribution, and retirements are open challenges that require ongoing attention.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.