Skip to main content
QUICK REVIEW

[Paper Review] Exploring Stereotypes and Biased Data with the Crowd

Zeyuan Hu, Julia Strout|arXiv (Cornell University)|Jan 10, 2018
Mobile Crowdsensing and Crowdsourcing18 references3 citations
TL;DR

This paper proposes using crowdsourcing via Amazon Mechanical Turk to detect hidden stereotypes and biases in machine learning training data, particularly beyond gender to include race and class. Despite challenges in data quality and worker ambiguity, the crowd identified diverse, non-obvious stereotypes—though many responses were redundant or low-quality—demonstrating modest but valuable potential for proactive bias detection in data collection workflows.

ABSTRACT

The goal of our research is to contribute information about how useful the crowd is at anticipating stereotypes that may be biasing a data set without a researcher's knowledge. The results of the crowd's prediction can potentially be used during data collection to help prevent the suspected stereotypes from introducing bias to the dataset. We conduct our research by asking the crowd on Amazon's Mechanical Turk (AMT) to complete two similar Human Intelligence Tasks (HITs) by suggesting stereotypes relating to their personal experience. Our analysis of these responses focuses on determining the level of diversity in the workers' suggestions and their demographics. Through this process we begin a discussion on how useful the crowd can be in tackling this difficult problem within machine learning data collection.

Motivation & Objective

  • To investigate whether crowdsourcing can proactively identify cultural stereotypes and biases in machine learning datasets before model training.
  • To extend beyond existing gender bias detection by exploring race, class, and intersectional stereotypes in data collection.
  • To evaluate the feasibility and quality of crowd-generated insights for detecting subtle, non-obvious biases in training data.
  • To inform data curation practices by identifying how diverse perspectives from crowd workers can improve dataset fairness.
  • To contribute to the broader effort of mitigating bias amplification in machine learning models through early detection in data collection.

Proposed method

  • Conducted two Human Intelligence Tasks (HITs) on Amazon Mechanical Turk, asking workers to suggest stereotypes related to race, class, and gender.
  • Collected open-ended responses from crowd workers based on personal experience, focusing on occupational, racial, and socioeconomic associations.
  • Analyzed responses for diversity, frequency, and semantic relevance, with manual filtering to identify rare or insightful stereotypes.
  • Used qualitative analysis to assess worker behavior, including adherence to instructions and potential misinterpretations of ambiguous tasks.
  • Evaluated the utility of image-based suggestions that subvert rather than reinforce stereotypes, proposing this as a future direction.
  • Implemented basic quality control by accepting all responses and performing manual inspection due to limitations in automated semantic validation.

Experimental results

Research questions

  • RQ1Can crowdsourcing effectively uncover non-obvious cultural stereotypes in training data that researchers might overlook?
  • RQ2To what extent do crowd workers identify biases beyond gender, such as those related to race and class, in data collection?
  • RQ3How does the diversity of crowd worker demographics contribute to detecting a broader range of stereotypes?
  • RQ4What are the challenges in ensuring quality and relevance of open-ended crowd responses for bias detection?
  • RQ5Can image-based suggestions that subvert stereotypes be a viable alternative to reinforcing them in bias detection tasks?

Key findings

  • The crowd successfully identified a range of non-gendered stereotypes, including racial, class-based, and intersectional associations such as 'manly looking transgender woman' and 'white rural American with Confederate flag'.
  • Despite the diversity of perspectives, 46.95% of verbs in the imSitu dataset were found to be biased toward one gender, with 47.5% of verbs showing mean amplification of 0.05% during model training.
  • MS-COCO data set showed one-third of noun objects heavily biased toward men, especially in sports, while kitchen objects were strongly female-biased, with some bias amplification up to 0.1%.
  • Many responses were redundant or low-quality, with workers frequently providing common or semantically irrelevant suggestions, such as 'I WANT A CLASS TYPE', highlighting data quality challenges.
  • A small number of workers offered creative, stereotype-subverting image suggestions, suggesting that future tasks could explicitly solicit such counter-stereotypical content to reduce repetition.
  • The study found no malicious or overtly offensive comments, but some workers appeared to misunderstand the task or were unengaged, indicating the need for better task design and quality control.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.