Skip to main content
QUICK REVIEW

[论文解读] Big Data Quality: A systematic literature review and future research directions

Mostafa Mirzaie, Behshid Behkamal|arXiv (Cornell University)|Apr 10, 2019
Data Quality and Management参考文献 84被引用 5
一句话总结

本文对2009年至2019年期间的大数据质量研究进行了系统性文献回顾,提出了一种类层次框架,按数据处理类型(流式、批处理、混合型)、主要任务和评估方法对研究进行分类。该研究识别出现有方法中的关键缺口,并为提升实际应用中大数据质量的未来研究方向提供了建议。

ABSTRACT

One of the most significant problems of Big Data is to extract knowledge through the huge amount of data. The usefulness of the extracted information depends strongly on data quality. In addition to the importance, data quality has recently been taken into consideration by the big data community and there is not any comprehensive review conducted in this area. Therefore, the purpose of this study is to review and present the state of the art on the quality of big data research through a hierarchical framework. The dimensions of the proposed framework cover various aspects in the quality assessment of Big Data including 1) the processing types of big data, i.e. stream, batch, and hybrid, 2) the main task, and 3) the method used to conduct the task. We compare and critically review all of the studies reported during the last ten years through our proposed framework to identify which of the available data quality assessment methods have been successfully adopted by the big data community. Finally, we provide a critical discussion on the limitations of existing methods and offer suggestions on potential valuable research directions that can be taken in future research in this domain.

研究动机与目标

  • 为应对大数据环境中数据质量保障日益严峻的挑战,因为低质量数据会削弱知识提取与决策制定能力。
  • 识别并分析过去十年中大数据质量研究的最新进展,重点关注方法论趋势与应用场景。
  • 评估现有数据质量评估方法在不同大数据处理类型与任务中的有效性与采用情况。
  • 指出当前方法的局限性,并为未来大数据质量研究提供可操作的、基于证据的建议。

提出的方法

  • 在主要学术数据库与资源库(包括arXiv和DBLP)中,使用预定义的搜索标准开展系统性文献回顾。
  • 开发一种类层次框架,基于三个维度对大数据质量研究进行分类:处理类型(流式、批处理、混合型)、主要任务(如数据清洗、验证)和评估方法。
  • 分析2009年至2019年间发表的126项相关研究,以描绘大数据质量研究中的趋势、方法与应用领域。
  • 对现有数据质量评估技术进行批判性比较,重点关注其适用性、可扩展性与实际应用场景。
  • 识别出在不同大数据处理流水线中数据质量评估存在的重复性方法论缺陷与不一致之处。
  • 将研究发现整合为结构化讨论,总结当前方法的局限性,并提出针对性的未来研究方向。

实验结果

研究问题

  • RQ1在2009年至2019年的大数据研究中,哪些数据质量评估方法被最频繁采用?
  • RQ2在流式、批处理与混合型大数据处理类型中,数据质量方法有何差异?
  • RQ3大数据质量研究中最常见的任务是什么?针对每类任务,哪种方法最为有效?
  • RQ4当前大数据环境中数据质量评估技术的主要局限性是什么?
  • RQ5在实践中推动大数据质量发展的未来研究方向中,哪些最具前景?

主要发现

  • 大多数大数据质量研究集中于批处理,而针对实时流式处理的研究显著较少,表明在动态数据质量管理方面存在研究空白。
  • 规则驱动验证、统计分析与机器学习等数据质量评估方法被广泛使用,但其与大数据平台的集成仍有限。
  • 仅有不到15%的研究在生产环境中评估数据质量,凸显了在真实部署中缺乏实证验证。
  • 大数据系统中缺乏标准化的度量指标与基准测试,导致评估实践不一致。
  • 尽管现代数据流水线中广泛使用混合处理模型,但其在数据质量研究中仍被严重忽视。
  • 未来研究应优先关注可扩展、自动化且具备上下文感知能力的数据质量技术,这些技术应能原生集成至Spark、Kafka等大数据处理框架中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。