Skip to main content
QUICK REVIEW

[Paper Review] DataPerf: Benchmarks for Data-Centric AI Development

Mark Mazumder, Colby Banbury|arXiv (Cornell University)|Jul 20, 2022
Machine Learning and Data Classification51 citations
TL;DR

DataPerf introduces a community-driven benchmark suite for evaluating data-centric AI and data-centric algorithms across multiple modalities, hosted on an online platform with extensible benchmarks and long-term maintenance. The first iteration covers speech and vision data selection, data cleaning, data acquisition, and prompting, with open-source baselines.

ABSTRACT

Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry.

Motivation & Objective

  • Shift ML benchmarking from models to data quality and data-centric development practices.
  • Provide a scalable, open platform for evaluating data-centric pipelines and datasets.
  • Foster community contributions with a working group and long-term stewardship.
  • Demonstrate practical data-centric tasks with real-world use cases across modalities.

Proposed method

  • Develop an online platform (Dynabench) integrated with MLCommons to host data-centric benchmarks.
  • Extend the platform to accept diverse submission artifacts (training subsets, containerized systems, etc.).
  • Define five initial benchmarks (speech data selection, vision data selection, debugging, data acquisition, adversarial Nibbler) with fixed model settings for fair data-centric comparison.
  • Provide baseline implementations and a public leaderboard to enable reproducibility and progress tracking.
  • Maintain DataPerf via a dedicated working group under MLCommons for ongoing benchmark development and sustainability.

Experimental results

Research questions

  • RQ1How can benchmarks be designed to evaluate data-centric improvements independently of model changes?
  • RQ2What data-centric techniques yield the greatest gains within fixed model architectures and budgets?
  • RQ3How can an online platform support diverse data-centric challenges at scale and with reproducible evaluation?
  • RQ4What real-world use cases best illustrate gains from data-centric AI across modalities?
  • RQ5How do data acquisition, cleaning, and selection strategies compare in effectiveness and cost?

Key findings

  • DataPerf provides an extensible, open-source platform (Dynabench) and a long-term governance model via MLCommons for sustainable data-centric benchmarking.
  • The initial suite covers diverse data-centric tasks—speech and vision data selection, debugging, data acquisition, and adversarial prompting—demonstrating the breadth of data-centric development beyond model optimization.
  • Baseline results and demonstrations show heterogeneity across data marketplaces and tasks, underscoring the value of careful data-centric strategy design.
  • Offline evaluation scripts and containerized submission artifacts reduce online compute demands and improve accessibility for participants.
  • A dedicated DataPerf Working Group coordinates ongoing benchmark development, community contributions, and platform maintenance, aiming for long-term impact in academia and industry.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.