Skip to main content
QUICK REVIEW

[Paper Review] Data Readiness for AI: A 360-Degree Survey

Kaveen Hiniduma, Suren Byna|arXiv (Cornell University)|Apr 8, 2024
Data Quality and ManagementDecision Sciences3 citations
TL;DR

This paper proposes a comprehensive taxonomy of Data Readiness for AI (DRAI) metrics by analyzing over 120 academic and industry sources, covering structured and unstructured data. It synthesizes existing metrics for data quality, fairness, privacy, and FAIR principles, offering a unified framework to standardize evaluation and improve AI model reliability through systematic data preparation.

ABSTRACT

Artificial Intelligence (AI) applications critically depend on data. Poor quality data produces inaccurate and ineffective AI models that may lead to incorrect or unsafe use. Evaluation of data readiness is a crucial step in improving the quality and appropriateness of data usage for AI. R&D efforts have been spent on improving data quality. However, standardized metrics for evaluating data readiness for use in AI training are still evolving. In this study, we perform a comprehensive survey of metrics used to verify data readiness for AI training. This survey examines more than 140 papers published by ACM Digital Library, IEEE Xplore, journals such as Nature, Springer, and Science Direct, and online articles published by prominent AI experts. This survey aims to propose a taxonomy of data readiness for AI (DRAI) metrics for structured and unstructured datasets. We anticipate that this taxonomy will lead to new standards for DRAI metrics that will be used for enhancing the quality, accuracy, and fairness of AI training and inference.

Motivation & Objective

  • To address the lack of standardized metrics for evaluating data readiness in AI training and inference.
  • To identify and categorize data readiness dimensions across structured and unstructured datasets from a broad literature base.
  • To integrate fairness, privacy, and bias-related metrics into data quality assessment for ethical AI development.
  • To propose a unified DRAI (Data Readiness for AI) taxonomy that supports consistent evaluation and benchmarking.
  • To guide practitioners and researchers in selecting appropriate metrics based on data type, application domain, and AI lifecycle stage.

Proposed method

  • Conducted a systematic survey of 120+ sources from ACM Digital Library, IEEE Xplore, reputable journals, and expert web articles.
  • Mapped data readiness dimensions into a 360-degree framework, integrating quality, accessibility, interoperability, and ethical dimensions.
  • Classified metrics into categories: completeness, consistency, accuracy, timeliness, fairness, privacy, and reusability (FAIR).
  • Analyzed existing tools such as IBM’s Data Quality Toolkit and FAIR-compliant data infrastructures like The Materials Data Facility.
  • Evaluated scoring mechanisms and threshold definitions for data readiness, emphasizing interpretability and context-dependency.
  • Proposed a holistic, evolving framework that supports dynamic, lifecycle-based assessment of data readiness across diverse AI applications.

Experimental results

Research questions

  • RQ1What are the core data readiness dimensions essential for effective AI training and inference across structured and unstructured data?
  • RQ2How do existing data quality metrics (e.g., completeness, accuracy) align with AI-specific requirements and ethical considerations?
  • RQ3To what extent are fairness, bias, and privacy metrics integrated into current data readiness evaluation frameworks?
  • RQ4What are the key challenges in defining universal thresholds for data readiness, and how can they be addressed contextually?
  • RQ5How can benchmark datasets and evaluation protocols be developed to enable meaningful comparison of DRAI metrics across tools and domains?

Key findings

  • The survey identified over 120 relevant sources, revealing a fragmented but growing ecosystem of data readiness metrics across AI and data science literature.
  • Core data quality dimensions—completeness, consistency, accuracy, and timeliness—remain foundational, but are insufficient without fairness and privacy integration.
  • FAIR principles (Findable, Accessible, Interoperable, Reusable) are increasingly recognized as essential for AI-ready data, especially in scientific and research domains.
  • There is no universal threshold for data readiness; acceptable levels are highly context-dependent and vary by application domain and data type.
  • Interpretability and visualization of DRAI metrics are critical for stakeholder adoption, yet remain underdeveloped in current tools.
  • The study highlights the need for dynamic, lifecycle-based assessment, as data readiness must be continuously monitored beyond initial preparation stages.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.