[论文解读] Energy Concerns with HPC Systems and Applications
本综述论文探讨了高性能计算(HPC)和嵌入式系统中的能效挑战,分析了硬件架构、能效指标、冷却系统及优化技术。研究表明,能效对降低运营成本和碳排放至关重要,关键发现显示,专用加速器(如TPU)和高效微控制器可使AI推理工作负载的性能/瓦特提升2–5倍,并在能效方面相比CPU实现20倍的提升。
For various reasons including those related to climate changes, {\em energy} has become a critical concern in all relevant activities and technical designs. For the specific case of computer activities, the problem is exacerbated with the emergence and pervasiveness of the so called {\em intelligent devices}. From the application side, we point out the special topic of {\em Artificial Intelligence}, who clearly needs an efficient computing support in order to succeed in its purpose of being a {\em ubiquitous assistant}. There are mainly two contexts where {\em energy} is one of the top priority concerns: {\em embedded computing} and {\em supercomputing}. For the former, power consumption is critical because the amount of energy that is available for the devices is limited. For the latter, the heat dissipated is a serious source of failure and the financial cost related to energy is likely to be a significant part of the maintenance budget. On a single computer, the problem is commonly considered through the electrical power consumption. This paper, written in the form of a survey, we depict the landscape of energy concerns in computer activities, both from the hardware and the software standpoints.
研究动机与目标
- 解决高性能计算(HPC)和嵌入式系统日益增长的能耗成本和碳足迹问题。
- 研究能效在HPC(因冷却和运营成本)和嵌入式系统(因电池容量有限)中作为关键约束的作用。
- 探究硬件、软件和系统级优化中能效感知设计的作用。
- 评估AI工作负载对能耗的影响,并探索降低其能耗足迹的优化策略。
- 全面概述适用于多样化计算平台的能效指标、性能分析工具、冷却技术和优化方法。
提出的方法
- 对HPC和嵌入式系统中能效指标、功耗测量及能效分析工具的现有文献进行综述。
- 对能效硬件进行分类与分析:GPU、TPU、FPGA、微控制器(如Arduino、Raspberry Pi、Coral Dev Board)及SoC。
- 评估CPU、GPU和微控制器的能效管理工具,包括运行时监控和动态电压/频率调节。
- 回顾冷却技术,特别是HPC系统中液体冷却的趋势,以管理散热并提升效率。
- 系统性分析软件栈和系统层级中静态、动态及混合能效优化技术。
- 研究AI特定的能效优化:量化、剪枝、模型压缩、神经架构搜索及知识蒸馏,其中硬件选择(如TPU、A100 GPU)是关键使能因素。

实验结果
研究问题
- RQ1现代HPC和嵌入式系统中,能耗与碳足迹指标之间如何相关?
- RQ2在AI推理和HPC工作负载中,哪些最有效的硬件加速器和微控制器平台可最大限度降低能耗?
- RQ3在不同系统层级中,动态与静态能效优化技术在性能和节能方面的表现如何比较?
- RQ4模型压缩和知识蒸馏在不牺牲准确率的前提下,能在多大程度上降低AI推理的能耗成本?
- RQ5冷却系统和系统级策略在实现可持续HPC与嵌入式计算中发挥何种作用?
主要发现
- Frontier百亿亿次系统功耗为21.1 MW,算力为1.102 EFlop/s;Fugaku系统功耗为29.9 MW,算力为442 PFlop/s;Lumi系统功耗为2.94 MW,算力为151.9 PFlop/s,凸显了HPC系统能耗成本的持续增长。
- 在AI推理工作负载中,Coral Dev Board微控制器在求解时间与能效方面均比Intel Skylake CPU提升20倍以上。
- 使用TPU或A100 GPU替代通用CPU,可使机器学习训练的性能/瓦特提升2至5倍。
- 量化和知识蒸馏可分别将模型大小减少40%和60%,同时保留BERT语言理解能力的97%。
- 静态功耗显著影响数据中心的能耗足迹,因此亟需有效的空闲状态功耗管理。
- 液体冷却和芯片级热管理正成为可扩展且高效的HPC冷却的关键解决方案,而量子冷却技术也正在探索用于未来系统。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。