Skip to main content
QUICK REVIEW

[Paper Review] Unsolved Problems in ML Safety

Dan Hendrycks, Nicholas Carlini|arXiv (Cornell University)|Sep 28, 2021
Adversarial Robustness in Machine Learning179 references88 citations
TL;DR

The paper outlines four core ML safety problems—robustness, monitoring, alignment, and systemic safety—and provides concrete research directions for each.

ABSTRACT

Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the technical problems that the field needs to address. We present four problems ready for research, namely withstanding hazards ("Robustness"), identifying hazards ("Monitoring"), reducing inherent model hazards ("Alignment"), and reducing systemic hazards ("Systemic Safety"). Throughout, we clarify each problem's motivation and provide concrete research directions.

Motivation & Objective

  • Motivate the need for proactive ML safety research to prevent costly failures.
  • Identify four key problem areas in ML safety: robustness, monitoring, alignment, and systemic safety.
  • Clarify motivations and provide concrete directions to start or continue research in each area.

Proposed method

  • Define four ML safety problem areas and articulate their motivations.
  • Survey existing challenges and propose broad research directions for each area.
  • Suggest benchmarks, architectures, and evaluation approaches to advance safety.
  • Discuss societal, regulatory, and emergent-capability considerations that influence alignment.

Experimental results

Research questions

  • RQ1What are the four principal unsolved problems in ML safety and why are they critical now?
  • RQ2What concrete research directions can advance robustness, monitoring, alignment, and systemic safety?
  • RQ3How can benchmarks, detectors, and evaluation strategies be developed to address safety hazards?
  • RQ4What are the societal and regulatory implications of deploying powerful ML systems?

Key findings

  • Proposes a four-problem safety roadmap: robustness, monitoring, alignment, and systemic safety.
  • Outlines actionable directions for each problem, including benchmarks, detectors, and evaluation methods.
  • Highlights emergent capabilities and hidden backdoors as central alignment and monitoring concerns.
  • Discusses the role of proactive safety research in shaping regulation and reducing deployment hazards.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.