[Paper Review] Gender Bias in Text: Labeled Datasets and Lexicons
This paper introduces a publicly available collection of labeled datasets and lexicons for detecting gender bias in English text, using an enhanced taxonomy that includes four bias subtypes: Generic He, Generic She, Explicit Marking of Sex, and Gendered Neologisms. The authors employ a hybrid approach combining automated data retrieval, human annotation with high inter-rater reliability (Krippendorff’s alpha = 0.75), and word embedding-based lexicon augmentation to create resources for supervised and unsupervised NLP models to detect and mitigate gender bias.
Language has a profound impact on our thoughts, perceptions, and conceptions of gender roles. Gender-inclusive language is, therefore, a key tool to promote social inclusion and contribute to achieving gender equality. Consequently, detecting and mitigating gender bias in texts is instrumental in halting its propagation and societal implications. However, there is a lack of gender bias datasets and lexicons for automating the detection of gender bias using supervised and unsupervised machine learning (ML) and natural language processing (NLP) techniques. Therefore, the main contribution of this work is to publicly provide labeled datasets and exhaustive lexicons by collecting, annotating, and augmenting relevant sentences to facilitate the detection of gender bias in English text. Towards this end, we present an updated version of our previously proposed taxonomy by re-formalizing its structure, adding a new bias type, and mapping each bias subtype to an appropriate detection methodology. The released datasets and lexicons span multiple bias subtypes including: Generic He, Generic She, Explicit Marking of Sex, and Gendered Neologisms. We leveraged the use of word embedding models to further augment the collected lexicons.
Motivation & Objective
- To address the lack of representative, labeled datasets and lexicons for detecting gender bias in English text.
- To improve upon an existing gender bias taxonomy by re-formalizing its structure, adding a new bias type (Gendered Neologisms), and mapping subtypes to automated detection methods.
- To collect, annotate, and augment a diverse set of representative sentences and terms to support machine learning and NLP-based detection of gender bias.
- To provide publicly accessible resources that enable automated detection and mitigation of gender bias in textual content.
- To enhance inter-rater reliability in bias annotation through structured guidelines and validation using Krippendorff’s alpha.
Proposed method
- Retrieved potentially biased sentences using information retrieval and filtering techniques targeting specific bias subtypes.
- Employed nine graduate-level annotators to label sentences based on a refined taxonomy and annotated examples, ensuring clarity and consistency.
- Computed Krippendorff’s alpha (0.75) to validate inter-rater reliability and ensure shared understanding of labeling criteria.
- Curated gendered neologisms from Urban Dictionary by filtering terms containing gender-exclusive substrings (e.g., 'man') and requiring at least 100 up-votes for community acceptance.
- Manually labeled the final 500 terms as exclusionary based on definitions, ensuring relevance and bias detection accuracy.
- Augmented lexicons using word embedding models to improve coverage and detection performance in NLP pipelines.
Experimental results
Research questions
- RQ1How can a comprehensive and structured taxonomy of gender bias subtypes be formalized to support automated detection in NLP?
- RQ2What is the inter-rater reliability of human annotation when labeling diverse forms of gender bias in text?
- RQ3How can gendered neologisms—newly coined, biased terms—be systematically identified and extracted from user-generated content?
- RQ4To what extent can word embedding models enhance the coverage and robustness of gender bias lexicons?
- RQ5What methodological pipeline ensures high-quality, representative, and publicly accessible datasets for gender bias detection?
Key findings
- The inter-rater reliability of the annotation process was confirmed with a Krippendorff’s alpha score of 0.75, indicating strong agreement and clarity in labeling guidelines.
- The final dataset includes 500 manually labeled gendered neologisms extracted from Urban Dictionary after filtering for gender-exclusive substrings and minimum community up-vote thresholds (≥100).
- The authors successfully collected and annotated representative sentences across four gender bias subtypes: Generic He, Generic She, Explicit Marking of Sex, and Gendered Neologisms.
- The use of word embedding models enabled effective augmentation of lexicons, improving their coverage for downstream NLP tasks.
- The study produced publicly available labeled datasets and exhaustive lexicons, addressing a critical gap in resources for gender bias detection in English text.
- The improved taxonomy now includes a new bias type—Gendered Neologisms—alongside restructured subtypes and mapped detection methodologies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.