Skip to main content
QUICK REVIEW

[Paper Review] Data Quality Assessment: Challenges and Opportunities

Sedir Mohammed, Ehrlinger, Lisa|arXiv (Cornell University)|Mar 1, 2024
Data Quality and Management4 citations
TL;DR

This paper proposes a systematic, five-facet framework for comprehensive data quality (DQ) assessment, integrating dimensions like accuracy, timeliness, and transparency to enable numeric scoring and context-aware evaluation. It addresses the lack of standardized DQ assessment by identifying challenges and technologies across data sources, systems, tasks, and human factors, aiming to support data cleaning, compliance, and AI model reliability.

ABSTRACT

Data-oriented applications, their users, and even the law require data of high quality. Research has divided the rather vague notion of data quality into various dimensions, such as accuracy, consistency, and reputation. To achieve the goal of high data quality, many tools and techniques exist to clean and otherwise improve data. Yet, systematic research on actually assessing data quality in its dimensions is largely absent, and with it, the ability to gauge the success of any data cleaning effort. We propose five facets as ingredients to assess data quality: data, source, system, task, and human. Tapping each facet for data quality assessment poses its own challenges. We show how overcoming these challenges helps data quality assessment for those data quality dimensions mentioned in Europe's AI Act. Our work concludes with a proposal for a comprehensive data quality assessment framework.

Motivation & Objective

  • To address the lack of systematic, comprehensive data quality (DQ) assessment methods despite extensive research on DQ dimensions.
  • To establish a unified framework that enables numeric, context-aware DQ scoring across diverse data types and use cases.
  • To identify and analyze challenges and opportunities in assessing DQ across five key facets: data source, system, task, human, and data itself.
  • To support data cleaning, regulatory compliance (e.g., GDPR, AI Act), and data-centric AI by enabling measurable DQ evaluation.
  • To bridge the gap between theoretical DQ dimensions and practical assessment by integrating technical, organizational, and human factors.

Proposed method

  • Proposes a five-facet framework for DQ assessment: data source, system, task, human, and data, each with distinct assessment challenges and requirements.
  • Integrates 29 representative DQ dimensions (e.g., accuracy, consistency, timeliness, traceability, transparency, understandability, uniqueness) into the assessment framework.
  • Emphasizes the need for metadata, provenance tracking, version control, and standardized documentation to enable reproducible and auditable DQ assessment.
  • Highlights the role of scalable data processing, automated profiling, and user-centric design in enabling large-scale, continuous DQ monitoring.
  • Proposes that DQ assessment must be context-sensitive, aligning with specific use cases, domain regulations (e.g., HIPAA, GDPR), and model training requirements.
  • Calls for integrated technologies across data lineage, data quality tools, and governance platforms to operationalize DQ assessment in practice.
Figure 1 . Facets of data quality assessment and the exemplary characteristics for various data quality dimensions.
Figure 1 . Facets of data quality assessment and the exemplary characteristics for various data quality dimensions.

Experimental results

Research questions

  • RQ1How can data quality be systematically assessed across multiple dimensions in a way that supports data cleaning and compliance?
  • RQ2What are the key challenges in assessing data quality across the five facets: source, system, task, human, and data?
  • RQ3How can data quality dimensions such as timeliness, traceability, and transparency be meaningfully measured and scored?
  • RQ4What role do domain-specific regulations (e.g., GDPR, HIPAA) and AI governance frameworks play in shaping DQ assessment?
  • RQ5How can a unified, numeric data quality profile be constructed that reflects the true quality of a dataset in its intended use context?

Key findings

  • The paper identifies five core facets—data source, system, task, human, and data—that form the foundation for a comprehensive data quality assessment framework.
  • It establishes that data quality cannot be meaningfully improved without systematic, numeric assessment, as emphasized by the principle that 'DQ cannot be improved if it cannot be measured'.
  • The assessment of dimensions like timeliness and traceability requires clear definitions of acceptable timeframes and provenance tracking mechanisms, respectively.
  • Transparency and understandability are critical for trust and compliance but pose challenges due to varying user backgrounds and the need for accessible, non-technical disclosure.
  • Uniqueness assessment is complicated by the need to define duplicate detection at multiple granularities (value, row, column, dataset) and the use of fuzzy matching techniques.
  • The paper highlights that regulatory frameworks (e.g., GDPR, AI Act) increase the complexity of DQ compliance due to potential contradictions and lack of standardized DQ dimension definitions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.