[Paper Review] Big Data Computing Using Cloud-Based Technologies, Challenges and Future Perspectives
This paper proposes a comprehensive survey of cloud-based big data computing, analyzing existing tools, frameworks, and mathematical techniques for processing large-scale data. It evaluates cloud adoption viability, identifies key challenges such as scalability and security, and outlines future research directions in distributed data analytics and infrastructure optimization.
The excessive amounts of data generated by devices and Internet-based sources at a regular basis constitute, big data. This data can be processed and analyzed to develop useful applications for specific domains. Several mathematical and data analytics techniques have found use in this sphere. This has given rise to the development of computing models and tools for big data computing. However, the storage and processing requirements are overwhelming for traditional systems and technologies. Therefore, there is a need for infrastructures that can adjust the storage and processing capability in accordance with the changing data dimensions. Cloud Computing serves as a potential solution to this problem. However, big data computing in the cloud has its own set of challenges and research issues. This chapter surveys the big data concept, discusses the mathematical and data analytics techniques that can be used for big data and gives taxonomy of the existing tools, frameworks and platforms available for different big data computing models. Besides this, it also evaluates the viability of cloud-based big data computing, examines existing challenges and opportunities, and provides future research directions in this field.
Motivation & Objective
- To analyze the current state of big data computing using cloud-based technologies.
- To identify and categorize existing tools, frameworks, and platforms for big data processing in cloud environments.
- To evaluate the viability of cloud computing for handling big data workloads.
- To examine major challenges in cloud-based big data systems, including performance, security, and scalability.
- To propose future research directions for advancing big data analytics in distributed and cloud-native architectures.
Proposed method
- Surveying and classifying existing big data computing tools and frameworks based on their architectural models and use cases.
- Analyzing mathematical and data analytics techniques applicable to big data, such as machine learning and statistical modeling.
- Evaluating cloud-based infrastructure capabilities in supporting dynamic scaling of storage and processing for big data workloads.
- Categorizing challenges in cloud-based big data systems into technical, operational, and security dimensions.
- Providing a taxonomy of big data computing models, including batch, stream, and real-time processing, with associated cloud-native solutions.
- Synthesizing insights from existing literature to identify gaps and opportunities for future research in cloud-native big data analytics.
Experimental results
Research questions
- RQ1What are the key cloud-based technologies and frameworks enabling scalable big data processing?
- RQ2How do mathematical and data analytics techniques enhance the value extraction from big data in cloud environments?
- RQ3What are the primary technical and operational challenges in deploying big data workloads on cloud platforms?
- RQ4How viable is cloud computing as a solution for the dynamic storage and processing demands of big data?
- RQ5What are the emerging research directions for improving performance, security, and efficiency in cloud-based big data systems?
Key findings
- Cloud computing provides a scalable and flexible infrastructure for handling the dynamic storage and processing demands of big data.
- A wide range of tools and frameworks—such as Hadoop, Spark, and Flink—have been developed to support various big data computing models in the cloud.
- Challenges in cloud-based big data computing include data security, privacy, latency, and efficient resource allocation across distributed environments.
- The integration of advanced analytics techniques like machine learning with cloud platforms enables real-time insights and decision-making from large-scale data.
- Future research should focus on optimizing cloud-native architectures for low-latency processing, enhancing data governance, and improving fault tolerance in distributed big data pipelines.
- There remains a significant need for standardized benchmarks and evaluation frameworks to compare the performance of different cloud-based big data solutions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.