[论文解读] Towards Modeling Energy Consumption of Xeon Phi
本文提出了一种建模框架,用于预测异构计算环境中英特尔至强融核协处理器的能耗、执行时间和功耗。通过在代理应用 CoMD 和 LULESH 上开展频率和强可扩展性实验,作者验证了一个能够区分计算密集型和延迟敏感型行为的模型,结果表明,对于这两种应用,对称执行模式下均未观察到能耗节省。
In the push for exascale computing, energy efficiency is of utmost concern. System architectures often adopt accelerators to hasten application execution at the cost of power. The Intel Xeon Phi co-processor is unique accelerator that offers application designers high degrees of parallelism, energy-efficient cores, and various execution modes. To explore the vast number of available configurations, a model must be developed to predict execution time, power, and energy for the CPU and Xeon Phi. An experimentation method has been developed which measures power for the CPU and Xeon Phi separately, as well as total system power. Execution time and performance are also captured for two experiments conducted in this work. The experiments, frequency scaling and strong scaling, will help validate the adopted model and assist in the development of a model which defines the host and Xeon Phi. The proxy applications investigated, representative of large-scale real-world applications, are Co-Design Molecular Dynamics (CoMD) and Livermore Unstructured Lagrangian Explicit Shock Hydrodynamics (LULESH). The frequency experiment discussed in this work is used to determine the time on-chip and off-chip to measure the compute- or latencyboundedness of the application. Energy savings were not obtained in symmetric mode for either application.
研究动机与目标
- 开发一种用于至强融核协处理器的能耗、功耗和执行时间的预测模型。
- 分析不同执行模式(特别是对称模式)对能效的影响。
- 通过片上和片外时间测量,将应用行为分类为计算密集型或延迟敏感型。
- 通过在代表性代理应用上进行受控实验,验证该模型。
- 通过支持合理的配置选择,为实现百亿亿次计算提供能效执行的指导。
提出的方法
- 开展频率可调实验,测量不同时钟频率下的性能和功耗。
- 执行强可扩展性实验,评估多个至强融核核心上的可扩展性和能效。
- 分别测量 CPU、至强融核和整个系统的功耗,以隔离加速器的能耗。
- 使用片上和片外时间指标,判断应用是计算密集型还是延迟敏感型。
- 将该模型应用于代表大规模科学工作负载的代理应用 Co-Design 分子动力学(CoMD)和 LULESH。
- 通过在异构系统上进行受控硬件测量获得的实测数据,验证该模型。
实验结果
研究问题
- RQ1时钟频率的变化如何影响至强融核上的能耗和性能?
- RQ2对于科学代理应用,至强融核上对称执行模式的能效如何?
- RQ3如何利用片上与片外时间比值将应用行为分类为计算密集型或延迟敏感型?
- RQ4可扩展性特征在多大程度上影响至强融核部署的能效?
- RQ5预测模型能否准确捕捉至强融核在 CPU+加速器异构配置下的能耗?
主要发现
- 在 CoMD 和 LULESH 中均未观察到对称执行模式下的能耗节省,表明该配置效率低下。
- 频率可调实验通过片上和片外时间分析,成功识别出计算密集型和延迟敏感型特征。
- 该模型通过分离 CPU 和至强融核组件的贡献,准确捕捉了功耗和能耗。
- 强可扩展性实验展示了可扩展性趋势,从而提升了模型的预测准确性。
- 代理应用 CoMD 和 LULESH 展现出截然不同的性能和能耗特征,验证了该模型对多样化工作负载的适应能力。
- 本研究证实,能耗建模对于优化百亿亿次计算工作负载中至强融核的配置至关重要。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。