Skip to main content
QUICK REVIEW

[Paper Review] Resilience of Deep Learning applications: a systematic literature review of analysis and hardening techniques

Cristiana Bolchini, Luca Cassano|arXiv (Cornell University)|Sep 27, 2023
Radiation Effects in ElectronicsEngineering3 citations
TL;DR

This systematic literature review analyzes 71 studies (2019–2023) on deep learning resilience against hardware faults, classifying fault models, error detection techniques, and hardening strategies. It identifies key research trends, tooling support, and open challenges, proposing a unified framework to guide future development of resilient DL systems in safety-critical applications.

ABSTRACT

Machine Learning (ML) is currently being exploited in numerous applications being one of the most effective Artificial Intelligence (AI) technologies, used in diverse fields, such as vision, autonomous systems, and alike. The trend motivated a significant amount of contributions to the analysis and design of ML applications against faults affecting the underlying hardware. The authors investigate the existing body of knowledge on Deep Learning (among ML techniques) resilience against hardware faults systematically through a thoughtful review in which the strengths and weaknesses of this literature stream are presented clearly and then future avenues of research are set out. The review is based on 220 scientific articles published between January 2019 and March 2024. The authors adopt a classifying framework to interpret and highlight research similarities and peculiarities, based on several parameters, starting from the main scope of the work, the adopted fault and error models, to their reproducibility. This framework allows for a comparison of the different solutions and the identification of possible synergies. Furthermore, suggestions concerning the future direction of research are proposed in the form of open challenges to be addressed.

Motivation & Objective

  • To map the current research landscape on deep learning resilience against hardware faults in safety-critical systems.
  • To classify existing methods and tools based on fault models, error models, and DL frameworks for improved discoverability.
  • To identify research gaps and synergies across 163 studies, focusing on transient and permanent hardware faults.
  • To propose open challenges and future research directions for building an ecosystem of resilient DL solutions.
  • To evaluate the maturity and impact of existing tools and frameworks through citation and co-authorship analysis.

Proposed method

  • Conducted a systematic literature review using PRISMA guidelines, filtering 163 papers and selecting 71 for in-depth analysis based on relevance and scope.
  • Developed a classification framework based on fault model (transient/permanent), error model, DL framework (e.g., PyTorch, TensorFlow), and hardening technique.
  • Analyzed co-authorship networks and publication venues using VOSviewer to identify research clusters and collaboration trends.
  • Evaluated citation impact and cross-references among studies to assess influence and knowledge flow in the field.
  • Collected and cataloged 13 open-source tools for fault injection and resilience evaluation, including TensorFI2, Ares, and Ranger.
  • Synthesized findings into a structured overview of research trends, tooling, and unresolved challenges in DL hardware fault resilience.

Experimental results

Research questions

  • RQ1What are the dominant fault and error models used in recent research on deep learning resilience?
  • RQ2How are fault injection and resilience evaluation techniques categorized across different deep learning frameworks?
  • RQ3What are the most frequently used tools and frameworks for analyzing and hardening DL models against hardware faults?
  • RQ4What are the key research clusters and collaborative networks in the field of DL resilience?
  • RQ5What are the major open challenges and future research directions in achieving hardware fault resilience in deep learning systems?

Key findings

  • The research community has significantly expanded its focus on hardware fault resilience in deep learning since 2019, with a steady increase in publications per year.
  • The most common fault models are transient and permanent faults in weights, activations, and parameters, with fault injection being the primary evaluation method.
  • A total of 13 open-source tools were identified, including TensorFI2, Ranger, and FIdelity, supporting fault injection and resilience analysis across multiple DL frameworks.
  • Co-authorship analysis revealed 14 distinct research clusters, with 68 authors contributing to at least three publications, indicating a mature and collaborative research ecosystem.
  • Publication venues such as IEEE Transactions on Dependable and Secure Computing and ACM Transactions on Embedded Computing are central to the dissemination of this research.
  • Despite progress, significant challenges remain in standardizing fault models, improving reproducibility, and building a unified ecosystem for resilience evaluation and hardening.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.