Skip to main content
QUICK REVIEW

[Paper Review] The Anatomy of Big Data Computing

Raghavendra Kune, Pramod Kumar Konugurthi|arXiv (Cornell University)|Sep 4, 2015
Big Data Technologies and Applications33 references4 citations
TL;DR

This paper presents a comprehensive taxonomy and layered architecture for Big Data Computing, integrating cloud infrastructures into Big Data Clouds to address scalability and performance challenges. It outlines core components, technologies, and workloads, while identifying open challenges in data curation, processing, and analytics across distributed environments, offering a foundational framework for future research and system design in large-scale data analytics.

ABSTRACT

Advances in information technology and its widespread growth in several areas of business, engineering, medical and scientific studies are resulting in information/data explosion. Knowledge discovery and decision making from such rapidly growing voluminous data is a challenging task in terms of data organization and processing, which is an emerging trend known as Big Data Computing; a new paradigm which combines large scale compute, new data intensive techniques and mathematical models to build data analytics. Big Data computing demands a huge storage and computing for data curation and processing that could be delivered from on-premise or clouds infrastructures. This paper discusses the evolution of Big Data computing, differences between traditional data warehousing and Big Data, taxonomy of Big Data computing and underpinning technologies, integrated platform of Big Data and Clouds known as Big Data Clouds, layered architecture and components of Big Data Cloud and finally discusses open technical challenges and future directions.

Motivation & Objective

  • To analyze the evolution and distinguishing characteristics of Big Data Computing compared to traditional data warehousing.
  • To define a taxonomy of Big Data computing, including data types, processing models, and workloads.
  • To propose an integrated platform combining Big Data and Cloud computing, termed Big Data Clouds.
  • To outline a layered architectural model for Big Data Clouds, detailing components and their interactions.
  • To identify open technical challenges and future research directions in scalable data processing and analytics.

Proposed method

  • The paper conducts a comparative analysis between traditional data warehousing and Big Data computing, highlighting differences in data volume, velocity, variety, and processing paradigms.
  • It proposes a taxonomy of Big Data computing based on data types (structured, semi-structured, unstructured), processing models (batch, stream, interactive), and workloads (analytics, machine learning, ETL).
  • The authors introduce the concept of Big Data Clouds as an integrated platform combining cloud infrastructure with Big Data technologies to enable elastic, scalable data processing.
  • A five-layered architecture for Big Data Clouds is defined: (1) Data Ingestion, (2) Data Storage, (3) Data Processing, (4) Data Analytics, and (5) Application Interface.
  • The paper evaluates underpinning technologies such as Hadoop, Spark, NoSQL databases, and cloud platforms (e.g., AWS, OpenStack), analyzing their roles in each architectural layer.
  • It synthesizes existing research and industry practices to identify persistent technical challenges in data curation, fault tolerance, and real-time processing.

Experimental results

Research questions

  • RQ1How does Big Data Computing differ fundamentally from traditional data warehousing in terms of data characteristics and processing requirements?
  • RQ2What are the key components and architectural layers that constitute a scalable Big Data Cloud platform?
  • RQ3Which core technologies and mathematical models enable efficient data curation and analytics in Big Data systems?
  • RQ4What are the major open challenges in achieving high-performance, fault-tolerant, and real-time data processing in distributed Big Data environments?
  • RQ5What future research directions are needed to advance the integration of Big Data and cloud computing?

Key findings

  • The paper establishes that Big Data computing is characterized by high volume, velocity, and variety of data, requiring new processing models beyond traditional data warehousing.
  • The proposed Big Data Clouds platform enables dynamic resource allocation and elastic scaling by integrating cloud infrastructure with Big Data frameworks.
  • The five-layered architecture effectively organizes Big Data system components, with clear separation of concerns from data ingestion to application exposure.
  • Hadoop and Spark are identified as foundational technologies for batch and real-time processing, respectively, while NoSQL databases support flexible schema handling.
  • Open challenges include efficient data curation, fault tolerance in distributed environments, and low-latency processing for streaming workloads.
  • The study concludes that future advancements must focus on optimizing data analytics pipelines, improving interoperability, and enhancing security and privacy in Big Data systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.