Skip to main content
QUICK REVIEW

[论文解读] A Data as a Service (DaaS) Model for GPU-based Data Analytics

John Olorunfemi Abe, Burak Berk Ustundaug|arXiv (Cornell University)|Feb 5, 2018
Cloud Computing and Resource Management参考文献 14被引用 4
一句话总结

本文提出了一种基于GPU加速的Data as a Service(DaaS)模型,用于实时大数据分析,利用基于GPU的并行处理和自组织映射(SOM)实现高效聚类。该模型在使用NVIDIA GPU进行预处理和聚类任务时表现出显著的性能提升,实现了更快、可扩展的分析,并提高了云环境和数据中心工作负载的SLA/QoS合规性。

ABSTRACT

Cloud-based services with resources to be provisioned for consumers are increasingly the norm, especially with respect to Big data, spatiotemporal data mining and application services that impose a user's agreed Quality of Service (QoS) rules or Service Level Agreement (SLA). Considering the pervasive nature of data centers and cloud system, there is a need for a real-time analytics of the systems considering cost, utility and energy. This work presents an overlay model of GPU system for Data As A Service (DaaS) to give a real-time data analysis of network data, customers, investors and users' data from the datacenters or cloud system. Using a modeled layer to define a learning protocol and system, we give a custom, profitable system for DaaS on GPU. The GPU-enabled pre-processing and initial operations of the clustering model analysis is promising as shown in the results. We examine the model on real-world data sets to model a big data set or spatiotemporal data mining services. We also produce results of our model with clustering, neural networks' Self-organizing feature maps (SOFM or SOM) to produce a distribution of the clustering for DaaS model. The experimental results thus far show a promising model that could enhance SLA and or QoS based DaaS.

研究动机与目标

  • 为应对云环境和数据中心环境中对实时、低成本、节能的大数据分析日益增长的需求。
  • 提升数据密集型云服务中的服务质量(QoS)和服务水平协议(SLA)合规性。
  • 设计一种可扩展的、面向中间件优化的DaaS架构,利用GPU并行处理支持机器学习工作负载。
  • 评估在真实时空和农业数据集上,使用GPU加速聚类的自组织映射(SOM)的有效性。
  • 证明基于GPU的系统在电子健康、网络安全和智慧城市建设中实现低延迟、高吞吐量数据分析的可行性。

提出的方法

  • 设计一种用于DaaS的叠加GPU系统模型,将GPU加速与机器学习流水线集成。
  • 采用支持CUDA的GPU架构(NVIDIA GF108GL 和 GF114)实现并行数据预处理和聚类。
  • 将自组织映射(SOM)作为核心聚类算法,用于建模数据分布并识别模式。
  • 实现中间件层,以管理GPU任务调度、跨虚拟集群的负载均衡和同步。
  • 应用资源分配约束(N个总节点,μ_i服务速率,λ_i到达速率)的优化模型,以最大化系统效用。
  • 使用对数-对数尺度图,将聚类系数与聚类ID进行对比,分析不同数据层级下的聚类性能。

实验结果

研究问题

  • RQ1如何有效利用基于GPU的并行处理来加速DaaS模型中的实时数据分析?
  • RQ2在GPU加速的DaaS系统中,使用自组织映射(SOM)在多大程度上提升了聚类的准确性和效率?
  • RQ3中间件层是否可以通过优化资源分配和负载均衡,提升基于GPU的云分析中SLA和QoS的合规性?
  • RQ4该模型在真实时空和农业数据集上的处理速度和聚类质量表现如何?
  • RQ5GPU硬件(如CUDA核心数、内存)对DaaS模型的可扩展性和性能有何影响?

主要发现

  • GPU加速的DaaS模型在数据预处理和聚类方面相比传统CPU方法显著提升了速度。
  • 自组织映射(SOM)有效实现了数据分布的可视化和聚类,展示了在GPU上强大的模式发现能力。
  • 该模型在NVIDIA GF108GL(96个CUDA核心)和GF114(384个CUDA核心)GPU上表现出更优的性能和可扩展性,处理速度有明显提升。
  • 随着聚类层级的增加,均方根误差(RMSE)值下降,表明模型准确性和收敛性得到改善。
  • 聚类系数与聚类ID的对数-对数图揭示了在不同数据层级下一致且明显的聚类行为。
  • 该模型在电子健康、网络安全、灾害管理及智慧城市建设等实时分析应用中展现出巨大潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。