Skip to main content
QUICK REVIEW

[Paper Review] A Survey on Data Quality Dimensions and Tools for Machine Learning

Yuhan Zhou, Fengjiao Tu|arXiv (Cornell University)|Jun 28, 2024
Data Quality and Management5 citations
TL;DR

The paper reviews 17 data quality evaluation/improvement tools from the last five years, defines four ML-focused data quality dimensions with twelve metrics, and offers a roadmap for open-source tool development and future trends like LLMs.

ABSTRACT

Machine learning (ML) technologies have become substantial in practically all aspects of our society, and data quality (DQ) is critical for the performance, fairness, robustness, safety, and scalability of ML models. With the large and complex data in data-centric AI, traditional methods like exploratory data analysis (EDA) and cross-validation (CV) face challenges, highlighting the importance of mastering DQ tools. In this survey, we review 17 DQ evaluation and improvement tools in the last 5 years. By introducing the DQ dimensions, metrics, and main functions embedded in these tools, we compare their strengths and limitations and propose a roadmap for developing open-source DQ tools for ML. Based on the discussions on the challenges and emerging trends, we further highlight the potential applications of large language models (LLMs) and generative AI in DQ evaluation and improvement for ML. We believe this comprehensive survey can enhance understanding of DQ in ML and could drive progress in data-centric AI. A complete list of the literature investigated in this survey is available on GitHub at: https://github.com/haihua0913/awesome-dq4ml.

Motivation & Objective

  • Define and consolidate four data quality dimensions for ML and map them to practical metrics.
  • Survey 17 open-source data quality tools developed in the last five years and compare their capabilities.
  • Analyze challenges in evaluating and improving data quality for ML and propose a development roadmap for tools.
  • Discuss emerging trends such as large language models and generative AI applications in data quality for ML.

Proposed method

  • Identify and synthesize DQ dimensions and metrics applicable to ML from existing literature.
  • Compile and categorize 17 open-source DQ evaluation/improvement tools and extract their core functions and metrics.
  • Perform a comparative analysis of tools across dimensions, metrics, and update timelines.
  • Propose a roadmap for designing open-source DQ tools tailored to ML and data-centric AI.
  • Discuss the role of LLMs and generative AI in DQ evaluation/improvement for ML.
Figure 1: Evolution of DQ evaluation/improvement tools across functions over time. The 6 core functions are data loading, data profiling, data integration, data transformation, automation and monitoring, and output and reports. Every tool supports the loading and output functions so the middle four
Figure 1: Evolution of DQ evaluation/improvement tools across functions over time. The 6 core functions are data loading, data profiling, data integration, data transformation, automation and monitoring, and output and reports. Every tool supports the loading and output functions so the middle four

Experimental results

Research questions

  • RQ1What are the core data quality dimensions and metrics most pertinent to machine learning workflows?
  • RQ2How do current open-source DQ tools compare in functionality, metrics, and ML-centric applicability?
  • RQ3What roadmap and design principles can guide future open-source DQ tools for ML and data-centric AI?
  • RQ4What is the impact of emerging AI technologies (e.g., LLMs) on data quality evaluation and improvement for ML?

Key findings

  • Identified four DQ dimensions (intrinsic, contextual, representational, accessibility) and twelve ML-relevant metrics.
  • Reviewed 17 data quality evaluation/improvement tools from the past five years and summarised their functions, metrics, and update histories.
  • Provided a comparative analysis highlighting tools’ strengths in profiling, monitoring, and ML-focused evaluation, and noted a trend toward automation and monitoring in 2024 updates.
  • Outlined a development roadmap for ML-oriented DQ tools, including framework design, functionality, and potential integration with data-centric AI practices.
  • Discussed emerging opportunities for large language models and generative AI to enhance DQ evaluation and improvement in ML contexts.
Figure 2: DQ dimensions, metrics, and corresponding tools. It showcases 4 dimensions and 12 DQ metrics in the first and second rows. Beneath each one, corresponding tools are listed, indicating their evaluation focus on the specific metrics and dimensions. The color of each tool represents the last
Figure 2: DQ dimensions, metrics, and corresponding tools. It showcases 4 dimensions and 12 DQ metrics in the first and second rows. Beneath each one, corresponding tools are listed, indicating their evaluation focus on the specific metrics and dimensions. The color of each tool represents the last

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.