Skip to main content
QUICK REVIEW

[论文解读] HPC Cloud for Scientific and Business Applications: Taxonomy, Vision, and Research Challenges

Marco A. S. Netto, Rodrigo N. Calheiros|arXiv (Cornell University)|Oct 24, 2017
Cloud Computing and Resource Management参考文献 110被引用 12
一句话总结

本文提出了高性能计算(HPC)云的全面分类法、愿景及研究议程,分析其在科学与商业工作负载中的作用。研究识别出性能隔离、资源分配、定价模式及混合云集成等方面的关键挑战,主张通过自适应资源管理与可持续定价策略,实现云环境中HPC工作负载的可行化。

ABSTRACT

High Performance Computing (HPC) clouds are becoming an alternative to on-premise clusters for executing scientific applications and business analytics services. Most research efforts in HPC cloud aim to understand the cost-benefit of moving resource-intensive applications from on-premise environments to public cloud platforms. Industry trends show hybrid environments are the natural path to get the best of the on-premise and cloud resources---steady (and sensitive) workloads can run on on-premise resources and peak demand can leverage remote resources in a pay-as-you-go manner. Nevertheless, there are plenty of questions to be answered in HPC cloud, which range from how to extract the best performance of an unknown underlying platform to what services are essential to make its usage easier. Moreover, the discussion on the right pricing and contractual models to fit small and large users is relevant for the sustainability of HPC clouds. This paper brings a survey and taxonomy of efforts in HPC cloud and a vision on what we believe is ahead of us, including a set of research challenges that, once tackled, can help advance businesses and scientific discoveries. This becomes particularly relevant due to the fast increasing wave of new HPC applications coming from big data and artificial intelligence.

研究动机与目标

  • 分析HPC云在科学与商业计算环境中的当前采用状况。
  • 识别基于云的HPC中性能隔离、资源分配及服务质量方面的关键研究挑战。
  • 通过不断演进的定价与合同模式,探讨HPC云的可持续性,以适应不同规模的用户。
  • 探索混合云架构在平衡本地部署与云工作负载方面的作用,以实现成本与性能的优化。
  • 评估新兴技术(如容器、FPGA、GPU及低延迟网络)对HPC云可行性的影响力。

提出的方法

  • 基于部署模式构建HPC云的分类法:'HPC in the cloud'、'HPC plus cloud' 和 'HPC as a Service'。
  • 利用HPC基准测试与真实应用场景(包括基于MPI的工作负载及极易并行化的工作负载)分析性能与成本的权衡。
  • 评估资源共享与虚拟化对公共云中应用性能波动的影响。
  • 提出HPC云的愿景:一种结合本地与云资源的混合、弹性且可编程的平台。
  • 研究新兴技术(如容器、FPGA、GPU及低延迟网络,例如Azure、ProfitBricks)作为HPC云的推动因素。
  • 识别在自适应资源分配、SLA感知调度及面向小型与大型用户的成本优化资源配置方面的研究需求。

实验结果

研究问题

  • RQ1在共享非InfiniBand网络的公共云环境中,如何使具有高处理器间通信需求的HPC应用实现可接受的性能?
  • RQ2哪些最有效的混合云策略可实现本地与云资源的平衡,以优化成本与性能?
  • RQ3如何设计可持续且灵活的定价模型,以满足HPC云中小型用户与企业用户的需求?
  • RQ4新兴技术(如GPU、FPGA及低延迟网络)在缩小本地集群与云平台之间性能差距方面发挥何种作用?
  • RQ5资源管理系统如何适应由云使用成本责任驱动的用户行为变化,尤其是在传统上‘免费’的HPC环境中?

主要发现

  • 极易并行化的工作负载在当前云基础设施上可实现良好性能,而紧密耦合、高通信量的HPC工作负载则因网络延迟而面临可扩展性限制。
  • 公共云中的资源共享导致即使在相同资源分配下,重复执行时性能仍存在显著波动。
  • 结合本地集群与云弹性扩展的混合云模型,为工作负载波动的组织提供了可持续且成本效益高的路径。
  • 向按需付费云模式的转变可能改变用户行为,减少资源过度分配,相较于传统HPC集群中缺乏成本可见性的情况。
  • 新兴技术(如GPU、FPGA及低延迟网络,例如Azure、ProfitBricks)预计可显著缩小本地与云HPC平台之间的性能差距。
  • HPC云的可持续采用不仅需要技术进步,还需在自适应定价模型与SLA驱动的资源分配机制方面开展研究。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。