Skip to main content
QUICK REVIEW

[Paper Review] SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

Paul Röttger, Fabio Pernisi|arXiv (Cornell University)|Apr 8, 2024
Topic Modeling4 citations
TL;DR

This paper presents a systematic review of 102 open datasets for evaluating and improving large language model (LLM) safety, identifying trends such as the rise of synthetic and specialized datasets, dominance of English-language resources, and underutilization of newer datasets in practice. The key contribution is a living, community-maintained catalogue at SafetyPrompts.com to standardize and improve LLM safety evaluation practices across research and industry.

ABSTRACT

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety. However, much of this work has happened in parallel, and with very different goals in mind, ranging from the mitigation of near-term risks around bias and toxic content generation to the assessment of longer-term catastrophic risk potential. This makes it difficult for researchers and practitioners to find the most relevant datasets for their use case, and to identify gaps in dataset coverage that future work may fill. To remedy these issues, we conduct a first systematic review of open datasets for evaluating and improving LLM safety. We review 144 datasets, which we identified through an iterative and community-driven process over the course of several months. We highlight patterns and trends, such as a trend towards fully synthetic datasets, as well as gaps in dataset coverage, such as a clear lack of non-English and naturalistic datasets. We also examine how LLM safety datasets are used in practice -- in LLM release publications and popular LLM benchmarks -- finding that current evaluation practices are highly idiosyncratic and make use of only a small fraction of available datasets. Our contributions are based on SafetyPrompts.com, a living catalogue of open datasets for LLM safety, which we plan to update continuously as the field of LLM safety develops.

Motivation & Objective

  • Address the challenge of fragmented and rapidly growing LLM safety dataset landscape, which hinders researchers in selecting relevant datasets for specific safety goals.
  • Identify key trends, gaps, and patterns in dataset creation, format, language coverage, and licensing to guide future dataset development.
  • Assess how open safety datasets are currently used in model release publications and popular LLM benchmarks to expose inconsistencies and underutilization.
  • Establish a living, continuously updated catalogue of open LLM safety datasets at SafetyPrompts.com to support standardization and improve evaluation practices.
  • Highlight the disconnect between the availability of modern safety datasets and their actual use in real-world model evaluations, calling for improved standardization.

Proposed method

  • Conducted a systematic, community-driven search across academic publications, GitHub, and Hugging Face to identify open datasets relevant to LLM safety between June 2018 and February 2024.
  • Applied strict inclusion criteria: text-only datasets, relevance to safety (e.g., bias, toxicity, alignment, jailbreaking), availability via GitHub or Hugging Face, and no restrictions on licensing or language.
  • Analyzed 102 datasets along key dimensions: purpose, creation method (e.g., synthetic, real-world), format, size, access, licensing, and publication history.
  • Mapped dataset usage in 125 model release publications and 10 popular LLM benchmarks to assess real-world adoption patterns.
  • Categorized datasets by safety focus (e.g., sociodemographic bias, toxic generation, power-seeking behaviors) and evaluated language diversity and synthetic data prevalence.
  • Maintained a dynamic, publicly accessible dataset catalogue at SafetyPrompts.com, updated continuously with new dataset inclusions and metadata.

Experimental results

Research questions

  • RQ1What are the dominant trends in the creation and design of open datasets for LLM safety, particularly regarding data modality, language, and synthetic vs. real-world sources?
  • RQ2How are open safety datasets currently used in model release publications and established LLM benchmarks, and to what extent do they reflect the full diversity of available datasets?
  • RQ3What are the major gaps in current dataset coverage, especially in terms of language diversity and alignment with emerging safety risks such as sycophancy or power-seeking behaviors?
  • RQ4Why is there a persistent disconnect between the availability of new, high-quality safety datasets and their adoption in mainstream evaluation practices?
  • RQ5How can the field standardize LLM safety evaluation practices by leveraging the full breadth of existing open datasets?

Key findings

  • The creation of open LLM safety datasets has grown at an unprecedented rate, with over half of the 102 datasets reviewed published in 2023 or later.
  • There is a strong trend toward fully synthetic datasets, especially for specialized safety evaluations such as red-teaming, jailbreaking, and alignment testing.
  • English dominates the dataset landscape, with a clear lack of non-English safety datasets, particularly from non-US institutions.
  • Despite the availability of modern datasets, model release publications and benchmarks predominantly rely on older datasets from 2021–2022, such as BOLD and ToxiGen, which no longer reflect current LLM behavior.
  • Evaluation practices in both model releases and benchmarks are highly idiosyncratic, with minimal overlap in dataset usage, indicating a lack of standardization.
  • The review identifies a significant opportunity to improve LLM safety evaluation by integrating newer, more contextually relevant datasets into mainstream benchmarking and release evaluations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.