Skip to main content
QUICK REVIEW

[Paper Review] "Unsex me here": Revisiting Sexism Detection Using Psychological Scales and Adversarial Samples.

Mattia Samory, Indira Sen|arXiv (Cornell University)|Apr 27, 2020
Hate Speech and Cyberbullying Detection35 references4 citations
TL;DR

This paper introduces a theory-driven codebook for classifying sexism using psychological scales and adversarial samples, evaluating detection methods across datasets. It reveals current models fail to generalize beyond overt sexism, highlighting the need for broader, theory-informed approaches to improve detection robustness and inclusivity in NLP systems.

ABSTRACT

To effectively tackle sexism online, research has focused on automated methods for detecting sexism. In this paper, we use items from psychological scales and adversarial sample generation to 1) provide a codebook for different types of sexism in theory-driven scales and in social media text; 2) test the performance of different sexism detection methods across multiple data sets; 3) provide an overview of strategies employed by humans to remove sexism through minimal changes. Results highlight that current methods seem inadequate in detecting all but the most blatant forms of sexism and do not generalize well to out-of-domain examples. By providing a scale-based codebook for sexism and insights into what makes a statement sexist, we hope to contribute to the development of better and broader models for sexism detection, including reflections on theory-driven approaches to data collection.

Motivation & Objective

  • To develop a comprehensive codebook mapping psychological scale items to real-world sexist expressions in social media, grounded in established theories of sexism.
  • To evaluate the performance of existing sexism detection models across diverse datasets, assessing their robustness and generalization capabilities.
  • To analyze minimal human interventions that remove sexism, identifying linguistic patterns and strategies used to neutralize bias with minimal text changes.
  • To identify gaps in current detection methods, particularly their inability to capture subtle or context-dependent forms of sexism.
  • To advocate for theory-driven data collection and model development to improve the scope and fairness of automated sexism detection systems.

Proposed method

  • Adapted items from established psychological scales measuring hostile and benevolent sexism to create a structured codebook for labeling sexist content in social media text.
  • Generated adversarial samples by making minimal, targeted linguistic changes to sexist statements to test model robustness and identify vulnerabilities.
  • Evaluated multiple sexism detection models on multiple datasets to compare performance across in-domain and out-of-domain examples.
  • Analyzed human-annotated edits to sexist statements to extract common strategies for de-sexism, such as rephrasing, removing gendered assumptions, or replacing stereotypical language.
  • Used the codebook to map theoretical constructs of sexism (e.g., gender role stereotypes, objectification) to linguistic features in real-world text.
  • Integrated insights from psychological theory into NLP data collection and model evaluation to enhance representational breadth and fairness.

Experimental results

Research questions

  • RQ1How do psychological scale items map to real-world expressions of sexism in social media text?
  • RQ2To what extent do current automated sexism detection models generalize across different datasets and contexts?
  • RQ3What minimal linguistic changes do humans apply to remove sexism, and what do these changes reveal about the nature of sexist language?
  • RQ4How do adversarial samples expose weaknesses in existing sexism detection models?
  • RQ5In what ways can theory-driven approaches improve the design and evaluation of sexism detection systems?

Key findings

  • Current sexism detection models struggle to detect subtle or contextually embedded forms of sexism, performing best only on overtly hostile or explicit statements.
  • Models show poor generalization to out-of-domain data, indicating limited robustness across different social media contexts or linguistic styles.
  • Adversarial samples revealed that small, semantically minor changes to sexist statements often evade detection, highlighting model fragility.
  • Human strategies for removing sexism commonly involve replacing gendered stereotypes, rephrasing power dynamics, or removing objectifying language, suggesting that such features are critical for detection.
  • The codebook derived from psychological scales successfully mapped theoretical constructs of sexism to linguistic patterns in real text, offering a replicable framework for data annotation.
  • There is a clear need for theory-informed data collection and model design to improve detection of diverse and nuanced forms of sexism in online content.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.