Skip to main content
QUICK REVIEW

[Paper Review] Multimodal datasets: misogyny, pornography, and malignant stereotypes

Abeba Birhane, Vinay Uday Prabhu|arXiv (Cornell University)|Oct 5, 2021
Gender, Feminism, and Media56 references150 citations
TL;DR

The paper audits the LAION-400M multimodal dataset and exposes explicit misogynistic, pornographic, and biased content, discussing broader harms and open questions for stakeholders.

ABSTRACT

We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns surrounding the dubious curation practices used to generate these datasets, the sordid quality of alt-text data available on the world wide web, the problematic content of the CommonCrawl dataset often used as a source for training large language models, and the entrenched biases in large-scale visio-linguistic models (such as OpenAI's CLIP model) trained on opaque datasets (WebImageText). In the backdrop of these specific calls of caution, we examine the recently released LAION-400M dataset, which is a CLIP-filtered dataset of Image-Alt-text pairs parsed from the Common-Crawl dataset. We found that the dataset contains, troublesome and explicit images and text pairs of rape, pornography, malign stereotypes, racist and ethnic slurs, and other extremely problematic content. We outline numerous implications, concerns and downstream harms regarding the current state of large scale datasets while raising open questions for various stakeholders including the AI community, regulators, policy makers and data subjects.

Motivation & Objective

  • Assess the content and biases in large-scale LAION-400M multimodal datasets built from Common Crawl and CLIP-filtered pipelines.
  • Highlight risks of misogynistic, pornographic, racist, and other harmful content in image-text pairs.
  • Critique current curation, filtering, and detoxification practices in large datasets used for vision-language models.
  • Discuss ethical, regulatory, and practical implications for data subjects, AI developers, and policymakers.

Proposed method

  • Qualitative and quantitative examination of LAION-400M content revealed via CLIP-based filtering and alt-text analysis.
  • Description of the dataset construction pipeline: crawling a vast WWW corpus, filtering by CLIP-based similarity, and selecting image-text pairs.
  • Empirical assessment of NSFW prevalence in retrieved results for several queries using text and image filters.
  • Discussion of known biases in CLIP and potential mis-associations in filtered data.
  • Critical reflection on asymmetries between data collection (crawling) and downstream detoxification efforts.

Experimental results

Research questions

  • RQ1What is the prevalence and nature of explicit and harmful content (misogyny, pornography, stereotypes) in LAION-400M?
  • RQ2How do the filtering and curation pipelines (e.g., CLIP-based similarity thresholds) affect downstream harms and biases?
  • RQ3What are the ethical, regulatory, and practical implications of releasing and using such large-scale visio-linguistic datasets?
  • RQ4What asymmetries exist between data collection/curation and detoxification efforts, and how do they impact model harms?
  • RQ5What open questions should stakeholders (researchers, policymakers, data subjects) address regarding dataset composition and use?

Key findings

  • The LAION-400M search-audit revealed NSFW and explicit imagery linked to seemingly benign queries (e.g., Desi, Nun, Latina).
  • A substantial fraction of matches for sensitive terms contained NSFW indicators, illustrating risks of bias and harmful associations in retrieval results.
  • CLIP-based filtering thresholds (e.g., cosine similarity 0.3) can fail to prevent harmful content from being included, due to model biases and corner cases.
  • There are significant asymmetries between the ease of crawling/creating huge datasets and the effort required for downstream detoxification and harm reduction.
  • The dataset curation process often lacks robust, joint image-text filtering and can propagate biases and stereotypes.
  • The emotional toll and potential trauma on researchers performing sensitivity-level data curation is non-trivial and often underappreciated.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.