Skip to main content
QUICK REVIEW

[论文解读] Scale-Out Processors & Energy Efficiency.

Pouya Esmaili-Dokht, Mohammad Bakhshalipour|arXiv (Cornell University)|Aug 14, 2018
Parallel Computing and Optimization Techniques参考文献 18被引用 8
一句话总结

本文通過針對數據中心工作負載重新評估可擴展處理器,改以每瓦性能(P3)為優化目標,而非每單位面積性能(PD),證明了PD最佳的模組配置亦能最大化能源效率。研究確認,無論在不同製程節點或元件功耗變異下,16核心、4-MB LLC模組仍為最佳配置,證實可擴展處理器設計中能源效率與性能密度具有一致性。

ABSTRACT

Scale-out workloads like media streaming or Web search serve millions of users and operate on a massive amount of data, and hence, require enormous computational power. As the number of users is increasing and the size of data is expanding, even more computational power is necessary for powering up such workloads. Data centers with thousands of servers are providing the computational power necessary for executing scale-out workloads. As operating data centers requires enormous capital outlay, it is important to optimize them to execute scale-out workloads efficiently. Server processors contribute significantly to the data center capital outlay, and hence, are a prime candidate for optimizations. While data centers are constrained with power, and power consumption is one of the major components contributing to the total cost of ownership (TCO), a recently-introduced scale-out design methodology optimizes server processors for data centers using performance per unit area. In this work, we use a more relevant performance-per-power metric as the optimization criterion for optimizing server processors and reevaluate the scale-out design methodology. Interestingly, we show that a scale-out processor that delivers the maximum performance per unit area, also delivers the highest performance per unit power.

研究动机与目标

  • 以能源效率(每瓦性能)為主要優化指標,重新評估可擴展處理器設計,取代以往以每單位面積性能為優化目標的做法。
  • 確認針對性能密度(PD)最佳化的模組配置是否亦能最大化能源效率(P3)於可擴展工作負載中。
  • 分析核心、快取與DRAM功耗變異對最佳模組配置之影響。
  • 評估DRAM能源消耗在數據中心環境中對處理器設計決策之影響。

提出的方法

  • 本研究使用週期精確模擬、分析模型與技術報告,評估14奈米製程下、95W功耗預算與280 mm²面積限制之處理器架構。
  • 以每瓦性能(P3)為主要指標,比較傳統、陣列式與可擴展處理器架構。
  • 透過將元件功耗值(核心、LLC、DRAM)由基線值之0.1×至10×變動,掃描以評估配置之穩健性,進而找出最佳模組配置。
  • 針對16核心、4-MB LLC模組在多種配置下進行評估,以確認其在不同功耗條件下的穩定性。
  • 分析涵蓋有序(OoO)與無序核心設計,最多支援六組單通道DDR4記憶體控制器。
  • 結果透過多個製程節點與DRAM參數進行驗證,以評估發現之普遍性。

实验结果

研究问题

  • RQ1針對每單位面積性能(PD)優化之處理器,是否亦能在可擴展工作負載中達成每瓦性能(P3)之最佳化?
  • RQ2核心動態與靜態功耗、LLC功耗及DRAM存取能量之變異,如何影響最佳模組配置?
  • RQ316核心、4-MB LLC模組配置是否在不同製程節點與系統參數下均具備穩健性?
  • RQ4將DRAM能量納入優化模型時,對最佳模組大小與快取大小之選擇有何影響?
  • RQ5在何種條件下最佳模組配置會改變?其對元件級功耗變異之敏感度為何?

主要发现

  • 16核心、4-MB LLC模組配置在每瓦性能(P3)與每單位面積性能(PD)兩者中均保持最佳,顯示可擴展處理器設計中能源效率與性能密度具有一致性。
  • 核心動態功耗增加10倍,亦不會改變最佳模組配置,顯示對核心功耗變異具備高度穩健性。
  • 僅當核心靜態功耗增加8倍或LLC功耗增加4.7倍時,最佳模組配置才會改變,顯示僅極端功耗變動才會影響設計選擇。
  • DRAM存取能量具有顯著影響:DRAM存取功耗增加10倍時,最佳配置會傾向採用更大快取,以減少記憶體存取頻率。
  • 最佳模組配置在廣泛的製程節點與DRAM參數範圍內均保持穩定,顯示16核心、4-MB LLC設計具備廣泛適用性。
  • 將DRAM能量納入優化模型後,確認記憶體存取效率是決定可擴展工作負載最佳處理器架構之關鍵因素。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。