Skip to main content
QUICK REVIEW

[论文解读] Large Scale Low Power Computing System - Status of Network Design in ExaNeSt and EuroExa Projects

Roberto Ammendola, A. Biagioni|arXiv (Cornell University)|Apr 11, 2018
Interconnection Networks and Systems参考文献 1被引用 5
一句话总结

本文介绍了ExaNet,一种由欧盟资助的ExaNeSt和EuroExa项目所开发的低功耗、基于FPGA的可扩展互连网络架构,专为百亿亿次(exascale)高性能计算(HPC)系统设计。该架构基于Xilinx Zynq UltraScale+ FPGA构建,实现了512字节数据包的亚微秒级单跳延迟和高达90%的带宽效率,展现出在超大规模、高能效超级计算领域的强大潜力。

ABSTRACT

The deployment of the next generation computing platform at ExaFlops scale requires to solve new technological challenges mainly related to the impressive number (up to 10^6) of compute elements required. This impacts on system power consumption, in terms of feasibility and costs, and on system scalability and computing efficiency. In this perspective analysis, exploration and evaluation of technologies characterized by low power, high efficiency and high degree of customization is strongly needed. Among the various European initiative targeting the design of ExaFlops system, ExaNeSt and EuroExa are EU-H2020 funded initiatives leveraging on high end MPSoC FPGAs. Last generation MPSoC FPGAs can be seen as non-mainstream but powerful HPC Exascale enabling components thanks to the integration of embedded multi-core, ARM-based low power CPUs and a huge number of hardware resources usable to co-design application oriented accelerators and to develop a low latency high bandwidth network architecture. In this paper we introduce ExaNet the FPGA-based, scalable, direct network architecture of ExaNeSt system. ExaNet allow us to explore different interconnection topologies, to evaluate advanced routing functions for congestion control and fault tolerance and to design specific hardware components for acceleration of collective operations. After a brief introduction of the motivations and goals of ExaNeSt and EuroExa projects, we will report on the status of network architecture design and its hardware/software testbed adding preliminary bandwidth and latency achievements.

研究动机与目标

  • 应对百亿亿次HPC系统中日益严峻的功耗与可扩展性挑战,这些系统可能需要高达100万个计算核心并消耗高达100 MW的功耗。
  • 设计一种低功耗、高吞吐量、低延迟的互连网络,适用于利用可定制FPGA的超大规模系统。
  • 通过硬件与软件栈的协同设计,实现基于ARM的百亿亿次级超级计算机中的高效通信。
  • 开发并验证一种可扩展的基于RDMA的网络架构,支持集体操作与故障容错机制。
  • 通过硬件测试平台(KARMA)和初步基准测试,展示ExaNet架构的可行性与性能表现。

提出的方法

  • 利用高端MPSoC FPGA(Xilinx Zynq UltraScale+),结合嵌入式ARM Cortex-A53处理器与可重构逻辑,实现系统级集成。
  • 设计ExaNet为模块化、直接连接的网络架构,采用定制IP模块实现路由、流量控制及集体操作加速。
  • 使用Trenz FPGA板卡构建硬件测试平台(KARMA),用于评估网络性能与功耗表现。
  • 在基准测试中采用用户空间内存映射I/O与直接硬件访问机制,绕过内核驱动,以最小化延迟。
  • 采用自检机制,结合流量生成器、数据消费者与性能计数器,测量带宽与延迟。
  • 对网络组件(APErouter、APElink、Aurora收发器)进行低功耗与高吞吐量优化,重点关注FIFO深度与虚拟通道设计。

实验结果

研究问题

  • RQ1基于FPGA的互连网络能否在百亿亿次HPC系统中实现亚微秒级延迟与高带宽?
  • RQ2在基于FPGA的系统中,自定义网络IP的功耗如何随片内与片间端口数量的变化而变化?
  • RQ3与标准内核I/O相比,硬件优化的用户空间通信在HPC网络中能将延迟降低多少?
  • RQ4在真实硬件条件下,针对小到中等大小的数据包,自定义网络IP(ExaNet)的带宽效率能达到何种水平?
  • RQ5自适应路由与集体操作加速在超大规模系统中如何提升网络可扩展性与故障容错能力?

主要发现

  • ExaNet网络架构实现了约0.46 μs的单跳往返延迟,单跳传输时间为0.46 μs,表明其具备亚微秒级的节点间通信能力。
  • APErouter的功耗为0.088 W,其中72%归因于片内与片间端口,ExaNet网络IP总功耗仅为0.136 W。
  • APElink在512字节数据包大小下实现了90%的带宽效率,接近因SFP+连接器限制而产生的理论10 Gbps上限。
  • APErouter在512字节数据包下实现76%的效率,协议开销为6.25%;当使用两个发送端口同时发送至同一目标时,效率提升至89.5%。
  • KARMA单板的总功耗为3.5 W,其中FPGA占2.822 W,表明该平台具有极高的能效表现。
  • 结果表明,基于FPGA的协同设计互连网络能够满足百亿亿次系统对低延迟与高带宽的严苛要求,同时保持极低的功耗水平。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。