[논문 리뷰] A Survey on Data Quality Dimensions and Tools for Machine Learning
해당 논문은 지난 five years의 17개 데이터 품질 평가/개선 도구를 검토하고, ML에 초점을 둔 네 가지 데이터 품질 차원을 정의하며 twelve 개의 지표를 제시하고, 오픈 소스 도구 개발 로드맵과 LLMs과 같은 미래 트렌드를 다룹니다.
Machine learning (ML) technologies have become substantial in practically all aspects of our society, and data quality (DQ) is critical for the performance, fairness, robustness, safety, and scalability of ML models. With the large and complex data in data-centric AI, traditional methods like exploratory data analysis (EDA) and cross-validation (CV) face challenges, highlighting the importance of mastering DQ tools. In this survey, we review 17 DQ evaluation and improvement tools in the last 5 years. By introducing the DQ dimensions, metrics, and main functions embedded in these tools, we compare their strengths and limitations and propose a roadmap for developing open-source DQ tools for ML. Based on the discussions on the challenges and emerging trends, we further highlight the potential applications of large language models (LLMs) and generative AI in DQ evaluation and improvement for ML. We believe this comprehensive survey can enhance understanding of DQ in ML and could drive progress in data-centric AI. A complete list of the literature investigated in this survey is available on GitHub at: https://github.com/haihua0913/awesome-dq4ml.
연구 동기 및 목표
- Define and consolidate four data quality dimensions for ML and map them to practical metrics.
- Survey 17 open-source data quality tools developed in the last five years and compare their capabilities.
- Analyze challenges in evaluating and improving data quality for ML and propose a development roadmap for tools.
- Discuss emerging trends such as large language models and generative AI applications in data quality for ML.
제안 방법
- Identify and synthesize DQ dimensions and metrics applicable to ML from existing literature.
- Compile and categorize 17 open-source DQ evaluation/improvement tools and extract their core functions and metrics.
- Perform a comparative analysis of tools across dimensions, metrics, and update timelines.
- Propose a roadmap for designing open-source DQ tools tailored to ML and data-centric AI.
- Discuss the role of LLMs and generative AI in DQ evaluation/improvement for ML.

실험 결과
연구 질문
- RQ1What are the core data quality dimensions and metrics most pertinent to machine learning workflows?
- RQ2How do current open-source DQ tools compare in functionality, metrics, and ML-centric applicability?
- RQ3What roadmap and design principles can guide future open-source DQ tools for ML and data-centric AI?
- RQ4What is the impact of emerging AI technologies (e.g., LLMs) on data quality evaluation and improvement for ML?
주요 결과
- Identified four DQ dimensions (intrinsic, contextual, representational, accessibility) and twelve ML-relevant metrics.
- Reviewed 17 data quality evaluation/improvement tools from the past five years and summarised their functions, metrics, and update histories.
- Provided a comparative analysis highlighting tools’ strengths in profiling, monitoring, and ML-focused evaluation, and noted a trend toward automation and monitoring in 2024 updates.
- Outlined a development roadmap for ML-oriented DQ tools, including framework design, functionality, and potential integration with data-centric AI practices.
- Discussed emerging opportunities for large language models and generative AI to enhance DQ evaluation and improvement in ML contexts.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.