[Paper Review] A Data as a Service (DaaS) Model for GPU-based Data Analytics
This paper proposes a GPU-accelerated Data as a Service (DaaS) model for real-time big data analytics, leveraging GPU-based parallel processing and Self-Organizing Maps (SOM) for efficient clustering. The model demonstrates significant performance gains in pre-processing and clustering tasks using NVIDIA GPUs, achieving faster, scalable analytics with improved SLA/QoS compliance for cloud and data center workloads.
Cloud-based services with resources to be provisioned for consumers are increasingly the norm, especially with respect to Big data, spatiotemporal data mining and application services that impose a user's agreed Quality of Service (QoS) rules or Service Level Agreement (SLA). Considering the pervasive nature of data centers and cloud system, there is a need for a real-time analytics of the systems considering cost, utility and energy. This work presents an overlay model of GPU system for Data As A Service (DaaS) to give a real-time data analysis of network data, customers, investors and users' data from the datacenters or cloud system. Using a modeled layer to define a learning protocol and system, we give a custom, profitable system for DaaS on GPU. The GPU-enabled pre-processing and initial operations of the clustering model analysis is promising as shown in the results. We examine the model on real-world data sets to model a big data set or spatiotemporal data mining services. We also produce results of our model with clustering, neural networks' Self-organizing feature maps (SOFM or SOM) to produce a distribution of the clustering for DaaS model. The experimental results thus far show a promising model that could enhance SLA and or QoS based DaaS.
Motivation & Objective
- To address the growing need for real-time, cost-effective, and energy-efficient big data analytics in cloud and data center environments.
- To enhance Quality of Service (QoS) and Service Level Agreement (SLA) compliance in data-intensive cloud services.
- To design a scalable, middleware-optimized DaaS architecture leveraging GPU parallelism for machine learning workloads.
- To evaluate the effectiveness of GPU-accelerated clustering using Self-Organizing Maps (SOM) on real-world spatiotemporal and agricultural data.
- To demonstrate the feasibility of using GPU-based systems for low-latency, high-throughput data analytics in e-health, cybersecurity, and smart cities.
Proposed method
- Designs an overlay GPU system model for DaaS that integrates GPU acceleration with machine learning pipelines.
- Employs CUDA-enabled GPU architectures (NVIDIA GF108GL and GF114) for parallel data pre-processing and clustering.
- Utilizes Self-Organizing Maps (SOM) as the core clustering algorithm to model data distribution and identify patterns.
- Implements a middleware layer to manage GPU task scheduling, load balancing, and synchronization across virtual clusters.
- Applies optimization models with constraints on resource allocation (N total nodes, μ_i service rates, λ_i arrival rates) to maximize system utility.
- Uses a logarithmic, log-log scale visualization of cluster coefficients against cluster ID to analyze clustering performance across data levels.
Experimental results
Research questions
- RQ1How can GPU-based parallelism be effectively leveraged to accelerate real-time data analytics in a DaaS model?
- RQ2To what extent does the use of Self-Organizing Maps (SOM) improve clustering accuracy and efficiency in GPU-accelerated DaaS systems?
- RQ3Can a middleware layer enhance SLA and QoS compliance in GPU-based cloud analytics by optimizing resource allocation and load balancing?
- RQ4How does the proposed model perform on real-world spatiotemporal and agricultural data sets in terms of processing speed and clustering quality?
- RQ5What is the impact of GPU hardware (e.g., CUDA cores, memory) on the scalability and performance of the DaaS model?
Key findings
- The GPU-accelerated DaaS model achieved significantly faster data pre-processing and clustering compared to traditional CPU-based approaches.
- Self-Organizing Maps (SOM) effectively visualized and clustered data distributions, demonstrating strong pattern discovery capabilities on GPU.
- The model showed improved performance and scalability on NVIDIA GF108GL (96 CUDA cores) and GF114 (384 CUDA cores) GPUs, with measurable gains in processing speed.
- Root-mean-square error (RMSE) values decreased with increasing cluster levels, indicating improved model accuracy and convergence.
- The log-log plot of cluster coefficient against cluster ID revealed consistent and distinct clustering behavior across different data levels.
- The model demonstrated strong potential for real-time analytics in e-health, cyber-security, disaster management, and smart city applications.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.