Skip to main content
QUICK REVIEW

[Paper Review] A Survey of Big Data Machine Learning Applications Optimization in Cloud Data Centers and Networks

Sanaa Hamid Mohamed, Taisir E. H. El-Gorashi|arXiv (Cornell University)|Jan 1, 2019
Cloud Computing and Resource Management600 references6 citations
TL;DR

This survey presents a comprehensive analysis of optimization techniques for big data machine learning workloads in cloud data centers and networks, categorizing efforts into application-level, networking-level, and data center-level optimizations. It identifies key challenges such as traffic congestion, energy consumption, and multi-tenancy inefficiencies, and evaluates solutions leveraging virtualization, SDN, NFV, and containerization to improve performance, fairness, and energy efficiency across distributed systems.

ABSTRACT

This survey article reviews the challenges associated with deploying and optimizing big data applications and machine learning algorithms in cloud data centers and networks. The MapReduce programming model and its widely-used open-source platform; Hadoop, are enabling the development of a large number of cloud-based services and big data applications. MapReduce and Hadoop thus introduce innovative, efficient, and accelerated intensive computations and analytics. These services usually utilize commodity clusters within geographically-distributed data centers and provide cost-effective and elastic solutions. However, the increasing traffic between and within the data centers that migrate, store, and process big data, is becoming a bottleneck that calls for enhanced infrastructures capable of reducing the congestion and power consumption. Moreover, enterprises with multiple tenants requesting various big data services are challenged by the need to optimize leasing their resources at reduced running costs and power consumption while avoiding under or over utilization. In this survey, we present a summary of the characteristics of various big data programming models and applications and provide a review of cloud computing infrastructures, and related technologies such as virtualization, and software-defined networking that increasingly support big data systems. Moreover, we provide a brief review of data centers topologies, routing protocols, and traffic characteristics, and emphasize the implications of big data on such cloud data centers and their supporting networks. Wide ranging efforts were devoted to optimize systems that handle big data in terms of various applications performance metrics and/or infrastructure energy efficiency. Finally, some insights and future research directions are provided.

Motivation & Objective

  • To identify and categorize optimization strategies for big data machine learning applications in cloud environments.
  • To analyze challenges in cloud data centers and networks, including traffic congestion, energy consumption, and resource under/over-utilization.
  • To evaluate the impact of emerging technologies like SDN, NFV, and containers on big data system performance and efficiency.
  • To provide a structured review of optimization techniques across application, networking, and data center layers.
  • To highlight research gaps and future directions for scalable, energy-efficient, and high-performance big data systems in cloud infrastructures.

Proposed method

  • Systematic review of existing literature on big data and machine learning optimization in cloud data centers and networks.
  • Classification of optimization studies into three categories: application-level, networking-level, and data center-level optimizations.
  • Analysis of key technologies including MapReduce, Hadoop, SDN, NFV, virtual machines, containers, and data center topologies.
  • Evaluation of performance metrics such as completion time, fairness, cost, profit, and energy consumption across diverse workloads.
  • Synthesis of simulation-based and experimental results from real-world prototypes and cloud testbeds.
  • Identification of trade-offs between performance, energy efficiency, and resource utilization in distributed and geo-distributed frameworks.

Experimental results

Research questions

  • RQ1How do big data workloads impact cloud data center and network performance, particularly in terms of traffic, latency, and resource utilization?
  • RQ2What are the key challenges in optimizing big data machine learning applications across different layers of the cloud stack—application, network, and data center?
  • RQ3How can emerging technologies like SDN, NFV, and containers improve energy efficiency and performance in big data systems?
  • RQ4What are the trade-offs between performance, cost, and energy consumption in multi-tenant and geo-distributed big data environments?
  • RQ5What are the open research challenges in achieving scalable, fair, and energy-efficient big data processing in cloud infrastructures?

Key findings

  • Energy efficiency and performance are often in conflict, with most providers favoring over-provisioning to meet SLAs rather than minimizing power consumption.
  • SDN and NFV enable dynamic, application-aware network and resource management, reducing job completion time and improving energy efficiency.
  • Multi-tenancy introduces fairness and isolation challenges due to shared network and I/O resources, requiring dynamic scheduling and pricing models.
  • Geo-distributed frameworks face high latency and data transfer costs, necessitating new routing and resource allocation strategies.
  • Heterogeneous clusters lead to task completion time imbalances, requiring accurate profiling and intelligent scheduling algorithms.
  • Containerization and virtualization improve resource utilization and agility, especially in dynamic and large-scale data center environments.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.