[Paper Review] Benchmark Data Repositories for Better Benchmarking
This paper proposes that benchmark data repositories should be designed as first-class research infrastructure to improve ML benchmarking practices. By embedding persistent identifiers, standardized metadata, and versioning, repositories can address issues like dataset overuse, poor reproducibility, and undervalued data contributions—thereby enhancing the rigor, fairness, and sustainability of machine learning research evaluation.
In machine learning research, it is common to evaluate algorithms via their performance on standard benchmark datasets. While a growing body of work establishes guidelines for -- and levies criticisms at -- data and benchmarking practices in machine learning, comparatively less attention has been paid to the data repositories where these datasets are stored, documented, and shared. In this paper, we analyze the landscape of these $ extit{benchmark data repositories}$ and the role they can play in improving benchmarking. This role includes addressing issues with both datasets themselves (e.g., representational harms, construct validity) and the manner in which evaluation is carried out using such datasets (e.g., overemphasis on a few datasets and metrics, lack of reproducibility). To this end, we identify and discuss a set of considerations surrounding the design and use of benchmark data repositories, with a focus on improving benchmarking practices in machine learning.
Motivation & Objective
- To address the underappreciation of data work in machine learning by positioning benchmark repositories as venues for recognizing datasets as intellectual contributions.
- To identify systemic flaws in current benchmarking practices, including overreliance on a few datasets, lack of reproducibility, and insufficient documentation.
- To propose concrete design principles for benchmark repositories that improve dataset discoverability, citation, versioning, and ethical compliance.
- To connect repository-level practices to broader concerns in data-centric AI, such as representational harms and construct validity.
- To establish a framework for best practices in benchmark repository design that supports the entire dataset lifecycle—from creation to reuse.
Proposed method
- Propose the term 'benchmark data repository' to distinguish repositories focused on evaluation datasets from general data repositories.
- Advocate for persistent identifiers (e.g., DOIs) to enable reliable dataset citation and attribution, reducing reliance on citing associated papers alone.
- Introduce the concept of 'connection metadata' to link datasets to their original sources, methods, and evaluation contexts, improving traceability and reproducibility.
- Recommend standardized metadata schemas aligned with community standards (e.g., DataCite) to ensure consistency and machine-readability across repositories.
- Propose versioning and provenance tracking to support reproducibility and detect dataset drift or distribution shifts over time.
- Highlight the role of repositories in enforcing licensing and consent compliance, especially for datasets with sensitive or PII-containing data.
Experimental results
Research questions
- RQ1How can benchmark data repositories be designed to better support reproducible and fair machine learning benchmarking?
- RQ2What role do persistent identifiers and standardized metadata play in improving dataset citation and recognition of data as scholarly contributions?
- RQ3In what ways do current benchmarking practices fail due to overreliance on a narrow set of datasets and metrics?
- RQ4How can benchmark repositories help detect and mitigate representational harms and construct validity issues in evaluation datasets?
- RQ5What institutional and technical mechanisms can repositories implement to ensure ethical data use and compliance with licensing and consent standards?
Key findings
- Persistent identifiers such as DOIs are critical for enabling proper citation of datasets and recognizing data creators as first-class contributors to scholarly research.
- The absence of standardized metadata and versioning in many repositories undermines reproducibility and makes it difficult to track dataset evolution or detect distribution shifts.
- Benchmark repositories can serve as a central infrastructure for enforcing best practices in data sharing, including licensing, consent, and ethical review, especially for sensitive data.
- Current benchmarking practices are often skewed toward a small number of widely used datasets, which can lead to overfitting and limited generalization, a problem that repositories can help mitigate through diversity-aware curation.
- Repositories that support connection metadata and provenance tracking significantly improve the transparency and auditability of benchmark evaluations.
- The integration of citation, versioning, and licensing features into repositories can incentivize higher-quality dataset curation and documentation across the ML research lifecycle.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.