[Paper Review] AI-Driven Frameworks for Enhancing Data Quality in Big Data Ecosystems: Error_Detection, Correction, and Metadata Integration
This PhD thesis proposes AI-powered frameworks to detect, correct, and integrate metadata for data quality in big data ecosystems, including new metrics, anomaly detection, and corrective modeling.
The widespread adoption of big data has ushered in a new era of data-driven decision-making, transforming numerous industries and sectors. However, the efficacy of these decisions hinges on the quality of the underlying data. Poor data quality can result in inaccurate analyses and deceptive conclusions. Managing the vast volume, velocity, and variety of data sources presents significant challenges, heightening the importance of addressing big data quality issues. While there has been increased attention from both academia and industry, current approaches often lack comprehensiveness and universality. They tend to focus on limited metrics, neglecting other dimensions of data quality. Moreover, existing methods are often context-specific, limiting their applicability across different domains. There is a clear need for intelligent, automated approaches leveraging artificial intelligence (AI) for advanced data quality corrections. To bridge these gaps, this Ph.D. thesis proposes a novel set of interconnected frameworks aimed at enhancing big data quality comprehensively. Firstly, we introduce new quality metrics and a weighted scoring system for precise data quality assessment. Secondly, we present a generic framework for detecting various quality anomalies using AI models. Thirdly, we propose an innovative framework for correcting detected anomalies through predictive modeling. Additionally, we address metadata quality enhancement within big data ecosystems. These frameworks are rigorously tested on diverse datasets, demonstrating their efficacy in improving big data quality. Finally, the thesis concludes with insights and suggestions for future research directions.
Motivation & Objective
- Address gaps in current big data quality approaches across multiple metrics and universality.
- Introduce new quality metrics and a weighted scoring system for precise assessment.
- Propose a generic AI-based framework for detecting quality anomalies.
- Develop a predictive-model-based framework for correcting detected anomalies.
- Tackle metadata quality enhancement within big data ecosystems.
Proposed method
- Develop new quality metrics and a weighted scoring model for data quality assessment.
- Design a generic AI-based anomaly detection framework for diverse quality issues.
- Develop a predictive-model-based correction framework for detected anomalies.
- Integrate metadata quality enhancement mechanisms into big data pipelines.
- Validate the frameworks on diverse datasets to demonstrate feasibility.
Experimental results
Research questions
- RQ1How can a universal set of quality metrics and a weighted scoring system improve big data quality assessment?
- RQ2Can AI-based anomaly detection reliably identify quality issues across varied data sources?
- RQ3How effective are predictive models in correcting detected quality anomalies?
- RQ4How can metadata quality be assessed and integrated within big data ecosystems?
- RQ5What are the practical guidelines and limitations for applying these AI-driven frameworks?
Key findings
- Introduces a new set of data quality metrics and a weighted scoring system.
- Proposes a generic AI-based anomaly detection framework applicable across domains.
- Proposes a predictive-model-based correction framework for detected anomalies.
- Addresses metadata quality enhancement within big data ecosystems.
- Demonstrates feasibility and potential effectiveness of the frameworks on diverse datasets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.