Skip to main content
QUICK REVIEW

[Paper Review] A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability

Xiaowei Huang, Daniel Kroening|arXiv (Cornell University)|Dec 18, 2018
Adversarial Robustness in Machine Learning176 references53 citations
TL;DR

A comprehensive survey of safety and trustworthiness in deep neural networks, covering verification, testing, adversarial attacks/defense, and interpretability with 202 papers reviewed.

ABSTRACT

In the past few years, significant progress has been made on deep neural networks (DNNs) in achieving human-level performance on several long-standing tasks. With the broader deployment of DNNs on various applications, the concerns over their safety and trustworthiness have been raised in public, especially after the widely reported fatal incidents involving self-driving cars. Research to address these concerns is particularly active, with a significant number of papers released in the past few years. This survey paper conducts a review of the current research effort into making DNNs safe and trustworthy, by focusing on four aspects: verification, testing, adversarial attack and defence, and interpretability. In total, we survey 202 papers, most of which were published after 2017.

Motivation & Objective

  • Explain the concept of trustworthiness for DNNs through certification and explanation processes.
  • Review verification and testing techniques for DNN safety and reliability.
  • Summarize adversarial attack methods and corresponding defenses.
  • Survey interpretability approaches to make DNN decisions more understandable.

Proposed method

  • Systematic literature review of 202 papers published mainly after 2017.
  • Classification of safety properties such as local robustness, output reachability, and Lipschitzian properties.
  • Organization of techniques into verification (deterministic guarantees, bounds, and statistical guarantees), testing (coverage criteria and test case generation), attack/defense, and interpretability.

Experimental results

Research questions

  • RQ1What properties define safety and trustworthiness for DNNs? (e.g., robustness, reachability)
  • RQ2How can verification, testing, adversarial defense, and interpretability contribute to a certification and explanation framework for DNNs?
  • RQ3What are the main methodologies and guarantees offered by verification and testing approaches?
  • RQ4What are effective defense strategies against adversarial attacks and how are they certified?
  • RQ5What interpretability techniques help satisfy explanation requirements for trustworthy DNNs?

Key findings

  • DNN verification provides provable guarantees but struggles with scalability to large models.
  • Testing offers computationally lighter assurance with coverage-guided test case generation.
  • Adversarial attack techniques highlight vulnerabilities, while defenses aim to improve robustness and provide certified assurances.
  • Interpretability methods yield instance-wise and model-level explanations, improving user trust.
  • The survey emphasizes certification (pre-deployment) and explanation (lifetime) as core trust-building processes.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.