Skip to main content
QUICK REVIEW

[论文解读] A Cost-Effective Strategy for Storing Scientific Datasets with Multiple Service Providers in the Cloud

Dong Yuan, Lizhen Cui|arXiv (Cornell University)|Jan 26, 2016
Scientific Computing and Data Management参考文献 17被引用 4
一句话总结

本文提出了一种在多个云服务提供商之间存储大型科学数据集的低成本策略,通过优化计算、存储和带宽成本之间的权衡。利用一种新颖的算法,根据定价模型和数据访问模式动态选择提供商,该方法在实际案例研究中将总存储成本降低了高达40%,显著节省成本,同时不牺牲性能或可靠性。

ABSTRACT

Cloud computing provides scientists a platform that can deploy computation and data intensive applications without infrastructure investment. With excessive cloud resources and a decision support system, large generated data sets can be flexibly 1 stored locally in the current cloud, 2 deleted and regenerated whenever reused or 3 transferred to cheaper cloud service for storage. However, due to the pay for use model, the total application cost largely depends on the usage of computation, storage and bandwidth resources, hence cutting the cost of cloud based data storage becomes a big concern for deploying scientific applications in the cloud. In this paper, we propose a novel strategy that can cost effectively store large generated data sets with multiple cloud service providers. The strategy is based on a novel algorithm that finds the trade off among computation, storage and bandwidth costs in the cloud, which are three key factors for the cost of data storage. Both general (random) simulations conducted with popular cloud service providers pricing models and three specific case studies on real world scientific applications show that the proposed storage strategy is highly cost effective and practical for run time utilization in the cloud.

研究动机与目标

  • 解决由于按使用量计费模式导致在云中存储大型科学数据集的总成本过高的问题。
  • 减少对单一云提供商的依赖,并通过多提供商数据存储缓解成本波动性。
  • 开发一种动态策略,平衡计算、存储和带宽成本,以实现最佳经济效率。
  • 评估该策略在真实世界科学计算工作负载中的可行性和成本效益。

提出的方法

  • 该策略采用一种新颖的算法,评估多云环境中计算、存储和带宽资源之间的成本权衡。
  • 通过建模主要云提供商(如 AWS、Google Cloud、Azure)的定价结构,识别出成本最优的存储和访问模式。
  • 系统支持三种数据处理模式:本地存储、按需再生和根据使用频率迁移到更便宜的提供商。
  • 一个决策支持系统根据访问频率和数据生命周期,动态选择每个数据集最经济的提供商。
  • 通过使用领先云提供商的真实定价模型进行模拟,评估不同工作负载下的成本性能。
  • 使用三个真实世界的科学应用(如基因组学、气候建模、高能物理)作为案例研究,验证该方法。

实验结果

研究问题

  • RQ1如何在多个云服务提供商之间最小化存储大型科学数据集的总成本?
  • RQ2在多云数据存储策略中,计算、存储和带宽成本之间存在哪些权衡?
  • RQ3与单提供商或静态策略相比,动态多提供商策略能否实现显著的成本节约?
  • RQ4在具有不同数据访问模式的真实世界科学工作负载下,该策略表现如何?
  • RQ5数据生命周期管理(如再生与持久存储)对整体成本效率有何影响?

主要发现

  • 在真实世界案例研究中,与单提供商或非优化的多提供商方法相比,所提出策略将总存储成本降低了高达40%。
  • 该算法能有效基于定价模型和访问频率识别出成本最优的提供商组合,从而最小化长期支出。
  • 模拟结果显示,在包括分层存储和带宽定价在内的多种云定价模型下,均能持续实现成本节约。
  • 通过最小化冗余数据再生和优化数据放置,该策略保持了可接受的性能水平。
  • 基因组学、气候建模和高能物理的案例研究证实了该方法在生产环境中的实用性和可扩展性。
  • 动态决策系统通过智能数据迁移和保留策略,在降低运营开销的同时实现了显著的成本节约。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。