[Paper Review] NLPositionality: Characterizing Design Biases of Datasets and Models
NLPositionality introduces a framework to quantify design biases in NLP datasets and models by measuring alignment between annotations from a diverse global pool of volunteers and dataset labels or model predictions. The study reveals that datasets and models are most aligned with Western, White, college-educated, and younger populations, while non-binary individuals and non-native English speakers are systematically marginalized.
Design biases in NLP systems, such as performance differences for different populations, often stem from their creator's positionality, i.e., views and lived experiences shaped by identity and background. Despite the prevalence and risks of design biases, they are hard to quantify because researcher, system, and dataset positionality is often unobserved. We introduce NLPositionality, a framework for characterizing design biases and quantifying the positionality of NLP datasets and models. Our framework continuously collects annotations from a diverse pool of volunteer participants on LabintheWild, and statistically quantifies alignment with dataset labels and model predictions. We apply NLPositionality to existing datasets and models for two tasks -- social acceptability and hate speech detection. To date, we have collected 16,299 annotations in over a year for 600 instances from 1,096 annotators across 87 countries. We find that datasets and models align predominantly with Western, White, college-educated, and younger populations. Additionally, certain groups, such as non-binary people and non-native English speakers, are further marginalized by datasets and models as they rank least in alignment across all tasks. Finally, we draw from prior literature to discuss how researchers can examine their own positionality and that of their datasets and models, opening the door for more inclusive NLP systems.
Motivation & Objective
- To address the lack of tools for quantifying design biases in NLP systems stemming from researcher, dataset, and model positionality.
- To develop a scalable, continuous, and inclusive method for measuring how well NLP systems align with diverse populations.
- To expose and characterize systemic biases in existing datasets and models, especially for social acceptability and hate speech detection tasks.
- To promote awareness of positionality in NLP research and encourage researchers to examine their own and their systems’ positional backgrounds.
Proposed method
- Leveraging the LabintheWild platform to recruit 1,096 diverse annotators from 87 countries for continuous annotation collection.
- Collecting annotations using standardized Likert-scale ratings for social acceptability and toxicity, aligned with original dataset or model labels.
- Statistically quantifying alignment between annotator demographics and dataset/model outputs using correlation and agreement metrics.
- Applying the framework post-hoc to existing datasets and models, including GPT-4, without requiring retraining or model access.
- Using demographic self-reports (e.g., nationality, education, age, gender identity) to analyze positionality effects across identity groups.
- Prioritizing participant motivation and learning over monetary compensation to improve data quality and enable long-term, sustainable data collection.
Experimental results
Research questions
- RQ1How do NLP datasets and models align with the demographic backgrounds of annotators from diverse global regions?
- RQ2To what extent do design biases in NLP systems reflect the positionality of their creators and original annotators?
- RQ3Which demographic groups experience the least alignment with dataset labels and model predictions, indicating systemic marginalization?
- RQ4Can a continuous, volunteer-driven annotation framework effectively capture evolving positionality in NLP systems?
- RQ5How does the positionality of models like GPT-4 compare to that of fine-tuned models and datasets?
Key findings
- A total of 16,299 annotations were collected from 1,096 annotators across 87 countries, with an average of 38 annotations per day over more than a year.
- Datasets and models show the highest alignment with Western, White, college-educated, and younger populations—characteristic of WEIRD (Western, Educated, Industrialized, Rich, Democratic) demographics.
- Non-binary individuals and non-native English speakers exhibit the lowest alignment with dataset labels and model predictions across all tasks, indicating systemic marginalization.
- Original dataset annotators show strong alignment with their own labels, underscoring the role of creator positionality in shaping dataset biases.
- The framework successfully identifies positionality patterns in both fine-tuned models and general-purpose LLMs like GPT-4, revealing consistent alignment with privileged demographic groups.
- The study demonstrates that continuous, motivation-driven annotation on platforms like LabintheWild yields higher-quality data than paid crowdsourcing and enables long-term monitoring of design biases.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.