Skip to main content
QUICK REVIEW

[Paper Review] Interpreting Safety Outcomes: Waymo's Performance Evaluation in the Context of a Broader Determination of Safety Readiness

Francesca Favarò, Trent Victor|arXiv (Cornell University)|Jun 23, 2023
Risk and Safety AnalysisDecision Sciences3 citations
TL;DR

This paper critiques the reliance on aggregate crash data for evaluating Waymo's autonomous driving safety, arguing for a broader safety readiness assessment that integrates continuous in-use monitoring, event-level analysis, and alternative estimation techniques to address the 'credibility paradox' between ADS and human-driven baselines. It advocates for a diversified safety evaluation framework beyond simple outcome reporting.

ABSTRACT

This paper frames recent publications from Waymo within the broader context of the safety readiness determination for an Automated Driving System (ADS). Starting from a brief overview of safety performance outcomes reported by Waymo (i.e., contact events experienced during fully autonomous operations), this paper highlights the need for a diversified approach to safety determination that complements the analysis of observed safety outcomes with other estimation techniques. Our discussion highlights: the presentation of a "credibility paradox" within the comparison between ADS crash data and human-derived baselines; the recognition of continuous confidence growth through in-use monitoring; and the need to supplement any aggregate statistical analysis with appropriate event-level reasoning.

Motivation & Objective

  • To analyze the limitations of relying solely on aggregate crash data for assessing Waymo's autonomous driving system (ADS) safety performance.
  • To identify the 'credibility paradox' in comparing ADS crash rates with human-driven vehicle baselines.
  • To advocate for a broader safety readiness determination that incorporates continuous monitoring and event-level reasoning.
  • To propose complementary estimation techniques that enhance confidence in ADS safety beyond statistical outcomes.
  • To support a more holistic, evidence-based evaluation of automated driving system readiness.

Proposed method

  • The paper conducts a critical review of Waymo's published safety performance outcomes, focusing on contact events during fully autonomous operations.
  • It introduces the concept of a 'credibility paradox' where low crash rates in ADS may not be statistically credible when compared to human-driven baselines.
  • The authors emphasize continuous confidence growth through in-use monitoring, suggesting this should be a core component of safety assessment.
  • They advocate for supplementing aggregate statistics with detailed event-level analysis to understand causal factors behind safety outcomes.
  • The framework integrates multiple estimation techniques beyond observed outcomes to support safety readiness decisions.
  • The approach draws on principles from software engineering and safety-critical systems to evaluate ADS reliability and readiness.

Experimental results

Research questions

  • RQ1How credible are Waymo's reported safety outcomes when compared to human-driven vehicle performance?
  • RQ2Why does a low number of contact events in autonomous operations not necessarily imply high safety readiness?
  • RQ3What role does continuous in-use monitoring play in building confidence in ADS safety?
  • RQ4How can event-level analysis improve the interpretation of safety outcomes beyond aggregate statistics?
  • RQ5What alternative estimation techniques are needed to support a comprehensive safety readiness determination?

Key findings

  • The 'credibility paradox' arises when low crash rates in autonomous vehicles are questioned due to insufficient data volume, making them statistically less credible than human-driven baselines.
  • Continuous in-use monitoring enables ongoing confidence growth in ADS performance, which is essential for safety readiness evaluation.
  • Aggregate statistical analysis alone is insufficient for safety assessment; event-level reasoning is necessary to understand root causes of safety outcomes.
  • Supplementing observed outcomes with estimation techniques enhances the robustness of safety readiness determination.
  • A diversified safety evaluation framework is essential for credible, defensible assessments of automated driving systems.
  • The paper concludes that safety readiness cannot be determined by crash data alone, requiring integration of monitoring, analysis, and estimation methods.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.