Skip to main content
QUICK REVIEW

[論文レビュー] Large Scale Studies of Memory, Storage, and Network Failures in a Modern Data Center

Justin Meza|arXiv (Cornell University)|Jan 1, 2019
Advanced Data Storage Technologies被引用数 10
ひとこと要約

本学位論文は、FacebookのデータセンターにおけるDRAM、SSD、ネットワークデバイスの故障を、長期間にわたり大規模な実運用データを用いて実地にわたる調査・分析したものである。故障傾向をモデル化し、メモリエラーを67%削減するページオフライン化や物理的ページランダム化といった信頼性向上手法を提案・検証した。また、非単調なSSDの摩耗行動や、システム設計に不可欠なネットワーク故障パターンの解明も行った。

ABSTRACT

The workloads running in the modern data centers of large scale Internet service providers (such as Amazon, Baidu, Facebook, Google, and Microsoft) support billions of users and span globally distributed infrastructure. Yet, the devices used in modern data centers fail due to a variety of causes, from faulty components to bugs to misconfiguration. Faulty devices make operating large scale data centers challenging because the workloads running in modern data centers consist of interdependent programs distributed across many servers, so failures that are isolated to a single device can still have a widespread effect on a workload. In this dissertation, we measure and model the device failures in a large scale Internet service company, Facebook. We focus on three device types that form the foundation of Internet service data center infrastructure: DRAM for main memory, SSDs for persistent storage, and switches and backbone links for network connectivity. For each of these device types, we analyze long term device failure data broken down by important device attributes and operating conditions, such as age, vendor, and workload. We also build and release statistical models to examine the failure trends for the devices we analyze. Our key conclusion in this dissertation is that we can gain a deep understanding of why devices fail---and how to predict their failure---using measurement and modeling. We hope that the analysis, techniques, and models we present in this dissertation will enable the community to better measure, understand, and prepare for the hardware reliability challenges we face in the future.

研究の動機と目的

  • 大規模データセンターにおけるDRAM、SSD、ネットワークデバイスの実世界の故障特性を理解すること。
  • デバイスの年齢、ベンダー、ワークロード、運用環境の要因に基づく故障傾向をモデル化すること。
  • ページオフライン化や物理的ページランダム化といった信頼性向上技術を大規模に評価・導入すること。
  • ハードウェア障害がソフトウェアシステムおよび高可用性Webサービスに与える影響を定量化すること。
  • 将来のデータセンター信頼性向上に向けたデータ駆動型のモデルと設計インサイトを提供すること。

提案手法

  • Facebookのサーバーフレットで14か月間にわたるDRAMエラー記録を収集・分析し、数十億デバイス日分のデータを扱った。
  • 数千台のサーバークラスタでページオフライン化を大規模に実装・評価し、メモリエラーを低減した。
  • ほぼ4年間にわたるSSD運用データを分析し、書き込み/読み取り量、消去/コピー操作、温度、電力の変動を追跡した。
  • データセンター内ネットワークの障害事例を7年間、WANの修理チケットを18か月分分析し、信頼性の傾向を評価した。
  • デバイスの属性、ワークロード、環境要因に基づいて故障率の統計モデルを構築した。
  • 物理的ページランダム化がDRAM故障率を低減する手法として有効であるかを評価し、その性能オーバーヘッドを測定した。

実験結果

リサーチクエスチョン

  • RQ1実世界のデータセンターにおいて、DRAMエラー率はデバイスの年齢、ベンダー、システム構成にどのように影響を受けるか?
  • RQ2フラッシュベースのSSDにおける主な故障パターンは何か。また、摩耗や運用環境の変化に伴い、そのパターンはどのように変化するか?
  • RQ3データセンター内およびデータセンター間ネットワークの障害は、システムの可用性やソフトウェア設計にどのように影響を与えるか?
  • RQ4ページオフライン化および物理的ページランダム化は、生産環境においてどれほどメモリエラー率を低減できるか?
  • RQ5ハードウェア障害の特性は、大規模分散システムの設計と信頼性にどのように影響を与えるか?

主な発見

  • ページオフライン化により、数千台のサーバーにわたる実運用環境で、メモリエラー率が67%削減された。
  • 物理的ページランダム化は、許容できる性能オーバーヘッドの範囲で信頼性向上の可能性を示したが、実運用導入には実世界の課題が伴った。
  • システム設計の選択、たとえば低密度DIMMや1チップあたりのコア数を減らすことで、DRAM障害率を最大57.7%まで低減できる。
  • SSDの故障率は摩耗に伴い単調に増加するのではなく、明確な故障発生段階と検出段階を示す。
  • データセンター内およびデータセンター間ネットワークの障害は、サービスの可用性に顕著な影響を与え、システムレベルでの緩和策が不可欠であることが判明した。
  • 本研究では、故障パターンがデバイスの属性、ワークロード、環境条件に強く依存しており、正確な予測にはデータ駆動型モデリングが不可欠であることが明らかになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。