[论文解读] Understanding Fault Scenarios and Impacts through Fault Injection Experiments in Cielo
本文针对Cielo petaflop级Cray XE超级计算机开展了一项大规模故障注入实验,旨在研究互连网络与计算节点中的故障到失效的传播机制。通过HPCArrow工具,作者在链路、节点和刀片级别注入硬件级故障,揭示了关键的恢复延迟与网络死锁现象,从而支持主动弹性机制的实现,其发现可推广至未来的Cray XC(Aries)系统。
We present a set of fault injection experiments performed on the ACES (LANL/SNL) Cray XE supercomputer Cielo. We use this experimental campaign to improve the understanding of failure causes and propagation that we observed in the field failure data analysis of NCSA's Blue Waters. We use the data collected from the logs and from network performance counter data 1) to characterize the fault-error-failure sequence and recovery mechanisms in the Gemini network and in the Cray compute nodes, 2) to understand the impact of failures on the system and the user applications at different scale, and 3) to identify and recreate fault scenarios that induce unrecoverable failures, in order to create new tests for system and application design. The faults were injected through special input commands to bring down network links, directional connections, nodes, and blades. We present extensions that will be needed to apply our methodologies of injection and analysis to the Cray XC (Aries) systems.
研究动机与目标
- 理解大规模高性能计算系统中,特别是网络与计算组件中的故障到失效的传播路径。
- 识别关键失效场景——尤其是网络死锁与长时间恢复时间——这些场景会影响应用程序的弹性。
- 开发并验证一种基于软件的故障注入工具(HPCArrow),以在真实超级计算机上系统化、可重复地开展故障实验。
- 为系统和应用级监控提供可操作的建议,以提升故障检测与恢复能力。
- 通过方法论的可迁移性,将Cielo(Cray XE)的洞察扩展至未来的Cray XC(Aries)系统。
提出的方法
- 开发了HPCArrow,一种基于软件实现的故障注入(SWIFI)工具,用于在真实petaflop级高性能计算系统中,对硬件组件(链路、节点、刀片)进行硬编码故障注入。
- 执行了18次受控的故障注入实验,针对互连网络链路、方向性连接、计算节点和刀片中的永久性故障。
- 通过日志和网络性能计数器监控系统行为,以捕获故障到失效的序列及恢复动态。
- 分析错误日志与恢复时间线,识别关键失效状态,如网络死锁与长时间恢复周期。
- 利用硬件错误日志中观察到的异常作为关键系统状态的指示器,以实现实时通知。
- 通过识别Cray XC(Aries)系统与Cray XE(Gemini)系统在网络行为与恢复机制上的相似性,将方法论扩展至未来在Cray XC系统上的故障注入。
实验结果
研究问题
- RQ1在大规模高性能计算系统的Gemini互连网络与计算节点中,故障到失效的主要传播路径是什么?
- RQ2网络相关故障,特别是链路与节点故障,如何影响应用程序执行与系统恢复时间?
- RQ3哪些关键失效状态——如网络死锁——会导致系统长时间无法恢复或完全不可恢复?
- RQ4如何改进系统与应用级的监控机制,以实现实时检测与响应网络故障?
- RQ5Cray XE(Cielo)系统中的故障注入方法论与发现,能在多大程度上适用于Cray XC(Aries)系统?
主要发现
- HPCArrow成功在Cielo的8,944个节点系统中对54条链路、2个节点和4个刀片注入了故障,验证了其在真实世界故障注入中的有效性。
- 观察到导致恢复时间足够长的网络死锁,足以支持对应用程序与系统软件的主动通知。
- 实验期间收集的硬件错误日志中包含可作为关键网络状态实时指示器的异常。
- 某些故障组合(如链路的顺序故障)的恢复时间超过数分钟,为应用程序创造了长时间的脆弱窗口。
- Cielo的Gemini网络中故障到失效的行为与恢复模式与Blue Waters系统中观察到的情况高度相似,验证了发现的可迁移性。
- Cray XE(Gemini)与Cray XC(Aries)系统在互连网络架构与恢复机制上的相似性,表明故障注入方法论可扩展至新一代系统。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。