[Paper Review] Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
This paper investigates whether AI safety benchmarks genuinely measure safety progress or merely reflect general model capabilities and training compute. Through a meta-analysis of dozens of models and benchmarks, it finds that many safety metrics are highly correlated with capabilities and compute—raising concerns about 'safetywashing'—and proposes empirical, decorrelated metrics to enable true differential safety progress beyond scaling.
As artificial intelligence systems grow more powerful, there has been increasing interest in "AI safety" research to address emerging and future risks. However, the field of AI safety remains poorly defined and inconsistently measured, leading to confusion about how researchers can contribute. This lack of clarity is compounded by the unclear relationship between AI safety benchmarks and upstream general capabilities (e.g., general knowledge and reasoning). To address these issues, we conduct a comprehensive meta-analysis of AI safety benchmarks, empirically analyzing their correlation with general capabilities across dozens of models and providing a survey of existing directions in AI safety. Our findings reveal that many safety benchmarks highly correlate with both upstream model capabilities and training compute, potentially enabling "safetywashing"--where capability improvements are misrepresented as safety advancements. Based on these findings, we propose an empirical foundation for developing more meaningful safety metrics and define AI safety in a machine learning research context as a set of clearly delineated research goals that are empirically separable from generic capabilities advancements. In doing so, we aim to provide a more rigorous framework for AI safety research, advancing the science of safety evaluations and clarifying the path towards measurable progress.
Motivation & Objective
- To investigate whether widely used AI safety benchmarks actually measure safety improvements or are confounded by upstream general capabilities and training compute.
- To challenge the assumption that capability gains equate to safety progress, especially in alignment and robustness benchmarks.
- To provide an empirical foundation for distinguishing safety advancements from general capability scaling, reducing the risk of misrepresenting capability gains as safety gains.
- To critique alignment theory as a top-down, intuition-driven paradigm that may mislead research priorities and advocate for empirically grounded safety evaluation instead.
- To guide model developers and researchers in selecting or designing benchmarks that are genuinely decorrelated from general capabilities to enable measurable, differential safety progress.
Proposed method
- Conduct a comprehensive meta-analysis of 40+ language models across multiple safety benchmarks and capabilities benchmarks.
- Compute a 'capabilities score' by distilling performance across general benchmark tasks (e.g., MMLU, GSM8K) to capture upstream model capabilities.
- Measure correlations between safety benchmark scores and both the derived capabilities score and raw training compute (FLOPs).
- Classify safety benchmarks by their correlation strength with capabilities and compute to identify those prone to safetywashing.
- Analyze the relationship between alignment techniques (e.g., RLHF, refusal fine-tuning) and safety benchmark performance to assess entanglement with scale.
- Propose a framework for designing new benchmarks that are empirically decorrelated from general capabilities, enabling measurement of true differential safety progress.
Experimental results
Research questions
- RQ1To what extent are AI safety benchmarks correlated with general model capabilities and training compute?
- RQ2Which safety benchmark categories (e.g., alignment, adversarial robustness, bias) show high vs. low correlations with upstream capabilities?
- RQ3Can safetywashing be empirically detected through correlations between safety metrics and capabilities/compute?
- RQ4How do alignment-based techniques influence safety benchmark performance, and are they truly improving safety or just scaling capabilities?
- RQ5What criteria can be used to design safety benchmarks that are empirically separable from general capabilities and thus measure true safety progress?
Key findings
- Approximately half of the evaluated AI safety benchmarks are highly correlated with general model capabilities and training compute, indicating a strong risk of safetywashing.
- Safety benchmarks in alignment, scalable oversight, truthfulness, and static adversarial robustness show high correlations with capabilities and compute, suggesting they may not measure safety independently.
- In contrast, benchmarks for bias, dynamic adversarial robustness, and calibration show relatively low correlations with capabilities, indicating they may be more effective at measuring distinct safety properties.
- Sycophancy and weaponization risk benchmarks exhibit significant negative correlations with general capabilities, suggesting that more capable models may be less prone to these behaviors.
- The study finds that alignment theory, while influential, often leads to research directions that are empirically indistinguishable from general capability scaling, undermining efforts to achieve differential safety progress.
- The paper concludes that current safety benchmarks are not reliably measuring safety progress independent of scaling, and calls for a shift toward empirically grounded, decorrelated evaluation frameworks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.