[论文解读] Resilience of Deep Learning applications: a systematic literature review of analysis and hardening techniques
本篇系统性文献回顾分析了2019至2023年间71项关于深度学习对硬件故障容错能力的研究,对故障模型、错误检测技术及加固策略进行分类。研究识别出关键研究趋势、工具支持以及开放性挑战,并提出一个统一框架,以指导安全关键应用中弹性深度学习系统的未来发展。
Machine Learning (ML) is currently being exploited in numerous applications being one of the most effective Artificial Intelligence (AI) technologies, used in diverse fields, such as vision, autonomous systems, and alike. The trend motivated a significant amount of contributions to the analysis and design of ML applications against faults affecting the underlying hardware. The authors investigate the existing body of knowledge on Deep Learning (among ML techniques) resilience against hardware faults systematically through a thoughtful review in which the strengths and weaknesses of this literature stream are presented clearly and then future avenues of research are set out. The review is based on 220 scientific articles published between January 2019 and March 2024. The authors adopt a classifying framework to interpret and highlight research similarities and peculiarities, based on several parameters, starting from the main scope of the work, the adopted fault and error models, to their reproducibility. This framework allows for a comparison of the different solutions and the identification of possible synergies. Furthermore, suggestions concerning the future direction of research are proposed in the form of open challenges to be addressed.
研究动机与目标
- 映射当前在安全关键系统中深度学习对硬件故障容错能力的研究现状。
- 基于故障模型、错误模型及深度学习框架,对现有方法和工具进行分类,以提升可发现性。
- 通过分析163项研究,识别出研究空白与协同效应,重点关注瞬态与永久性硬件故障。
- 提出开放性挑战与未来研究方向,以构建弹性深度学习解决方案的生态系统。
- 通过引用分析与合作者网络分析,评估现有工具与框架的成熟度与影响力。
提出的方法
- 依据PRISMA指南开展系统性文献回顾,筛选出163篇论文,并基于相关性与研究范围从其中选取71篇进行深入分析。
- 基于故障模型(瞬态/永久性)、错误模型、深度学习框架(如PyTorch、TensorFlow)以及加固技术,构建分类框架。
- 利用VOSviewer分析合作者网络与出版物来源,识别研究集群与合作趋势。
- 通过引用影响分析与文献间引用关系分析,评估该领域内影响力与知识流动情况。
- 收集并整理了13个用于故障注入与容错评估的开源工具,包括TensorFI2、Ares与Ranger。
- 将研究发现综合为关于研究趋势、工具支持及深度学习硬件故障容错领域未解挑战的结构化概览。
实验结果
研究问题
- RQ1近期关于深度学习容错能力的研究中,主要采用哪些故障与错误模型?
- RQ2在不同深度学习框架中,故障注入与容错评估技术如何分类?
- RQ3在分析与加固深度学习模型以应对硬件故障方面,最常使用的工具与框架有哪些?
- RQ4深度学习容错领域中的主要研究集群与协作网络是什么?
- RQ5在实现深度学习系统硬件故障容错方面,存在哪些主要开放性挑战与未来研究方向?
主要发现
- 自2019年以来,研究社区对深度学习硬件故障容错能力的关注显著扩大,每年发表的文献数量持续稳步增长。
- 最常见的故障模型为权重、激活值与参数中的瞬态与永久性故障,故障注入是主要的评估方法。
- 共识别出13个开源工具,包括TensorFI2、Ranger与FIdelity,支持在多个深度学习框架中进行故障注入与容错分析。
- 合作者网络分析揭示了14个独立的研究集群,其中68位作者至少参与了三篇及以上出版物,表明该研究生态系统已趋于成熟并具有高度协作性。
- IEEE Transactions on Dependable and Secure Computing与ACM Transactions on Embedded Computing等出版物是该研究领域知识传播的核心平台。
- 尽管已取得进展,但在标准化故障模型、提升可复现性以及构建统一的容错评估与加固生态系统方面,仍存在显著挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。