Skip to main content
QUICK REVIEW

[Paper Review] Large image datasets: A pyrrhic win for computer vision?

Vinay Uday Prabhu, Abeba Birhane|arXiv (Cornell University)|Jun 24, 2020
Ethics and Social Impacts of AISocial Sciences81 references44 citations
TL;DR

The paper critically audits ImageNet and related large-scale vision datasets, revealing ethical transgressions in data sourcing, labeling, and privacy; it provides a quantitative census and proposes audit-driven remedies.

ABSTRACT

In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably pornographic images in datasets. Taking the ImageNet-ILSVRC-2012 dataset as an example, we perform a cross-sectional model-based quantitative census covering factors such as age, gender, NSFW content scoring, class-wise accuracy, human-cardinality-analysis, and the semanticity of the image class information in order to statistically investigate the extent and subtleties of ethical transgressions. We then use the census to help hand-curate a look-up-table of images in the ImageNet-ILSVRC-2012 dataset that fall into the categories of verifiably pornographic: shot in a non-consensual setting (up-skirt), beach voyeuristic, and exposed private parts. We survey the landscape of harm and threats both society broadly and individuals face due to uncritical and ill-considered dataset curation practices. We then propose possible courses of correction and critique the pros and cons of these. We have duly open-sourced all of the code and the census meta-datasets generated in this endeavor for the computer vision community to build on. By unveiling the severity of the threats, our hope is to motivate the constitution of mandatory Institutional Review Boards (IRB) for large scale dataset curation processes.

Motivation & Objective

  • Assess the ethical implications and societal harms of large-scale vision datasets like ImageNet-ILSVRC-2012.
  • Quantify data governance issues through a cross-sectional census of age, gender, NSFW content, class semantics, and accuracy.
  • Propose corrective strategies and governance mechanisms to mitigate harms in LSVD curation.
  • Open-source the census datasets and code to enable community audits and transparency.

Proposed method

  • Perform a cross-sectional model-based census on ImageNet-ILSVRC-2012 across age, gender, NSFW content, class accuracy, and semanticity.
  • Hand-curate a look-up table to identify verifiably pornographic or non-consensual images within ImageNet-ILSVRC-2012.
  • Audit and visualize harms such as privacy loss, reverse image search risks, and gender bias using pre-trained models (DEX, InsightFace, RetinaFace, ArcFace).
  • Compile a dataset audit card and accompanying meta-datasets (CSV assets) to support transparency and reproducibility.
  • Open-source the code and census datasets to enable broader auditing by the CV community.
(a) Class-wise counts of the offensive classes
(a) Class-wise counts of the offensive classes

Experimental results

Research questions

  • RQ1To what extent do large-scale vision datasets contain ethically problematic images (e.g., verifiably pornographic, non-consensual, or harmful labels)?
  • RQ2What are the downstream privacy, consent, and bias risks associated with ImageNet and related LSVDs?
  • RQ3How can a quantitative census inform procedures to mitigate harms in dataset curation?
  • RQ4What governance and auditing mechanisms (e.g., dataset audit cards) can improve transparency and ethics in LSVDs?

Key findings

  • ImageNet contains non-consensual and potentially exploitative images across several classes, with hand-curated findings identifying misogynistic and pornographic content.
  • A census of 57 metrics across 61 class parameters reveals distributions related to count, age, gender, NSFW scores, and semanticity that highlight bias and privacy concerns.
  • The study documents ethical transgressions in WordNet-based class taxonomy and labeling practices that propagate stereotypes and privacy risks.
  • Reverse image search and opaque datasets (e.g., JFT-300M, Open Images, Tiny Images) enable real-world identification and privacy harms.
  • The authors provide open-source datasets and tutorials to facilitate ongoing auditing and advocate for Institutional Review Boards (IRBs) for LSVD curation.
  • An audit card example for ImageNet summarizes the dataset’s known shortcomings and the associated risks.
(b) Samples from the class labelled n****r
(b) Samples from the class labelled n****r

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.