[Paper Review] Semantics derived automatically from language corpora necessarily contain human biases.
This paper demonstrates that word embeddings like GloVe, trained on vast web text corpora, automatically learn and replicate human-like semantic biases—such as those related to race, gender, and social roles—because these biases are embedded in the language itself. Using novel evaluation tools (WEAT and WEFAT), the study shows that statistical machine learning models inherit societal prejudices not through design, but via exposure to biased language, revealing that language corpora encode historic and cultural biases that are then captured by AI systems.
Artificial intelligence and machine learning are in a period of astounding growth. However, there are concerns that these technologies may be used, either with or without intention, to perpetuate the prejudice and unfairness that unfortunately characterizes many human institutions. Here we show for the first time that human-like semantic biases result from the application of standard machine learning to ordinary language---the same sort of language humans are exposed to every day. We replicate a spectrum of standard human biases as exposed by the Implicit Association Test and other well-known psychological studies. We replicate these using a widely used, purely statistical machine-learning model---namely, the GloVe word embedding---trained on a corpus of text from the Web. Our results indicate that language itself contains recoverable and accurate imprints of our historic biases, whether these are morally neutral as towards insects or flowers, problematic as towards race or gender, or even simply veridical, reflecting the status quo for the distribution of gender with respect to careers or first names. These regularities are captured by machine learning along with the rest of semantics. In addition to our empirical findings concerning language, we also contribute new methods for evaluating bias in text, the Word Embedding Association Test (WEAT) and the Word Embedding Factual Association Test (WEFAT). Our results have implications not only for AI and machine learning, but also for the fields of psychology, sociology, and human ethics, since they raise the possibility that mere exposure to everyday language can account for the biases we replicate here.
Motivation & Objective
- To investigate whether machine learning models trained on everyday language inherit human-like semantic biases.
- To examine whether widely used NLP models such as GloVe reflect known psychological biases from studies like the Implicit Association Test.
- To develop new evaluation methods for detecting bias in word embeddings.
- To demonstrate that biases in language corpora are sufficient to produce biased semantic representations in AI models.
Proposed method
- Training a GloVe word embedding model on a large-scale web text corpus to learn dense vector representations of words.
- Applying the Word Embedding Association Test (WEAT) to measure associations between word categories (e.g., race, gender) and attributes (e.g., pleasant/unpleasant).
- Using the Word Embedding Factual Association Test (WEFAT) to evaluate associations between words and factual societal distributions (e.g., gender and careers, names and gender).
- Replicating known psychological biases from the Implicit Association Test using statistical associations in word embeddings.
- Comparing model-derived associations to human-observed biases to validate the replication of semantic biases.
- Analyzing the consistency and accuracy of bias patterns across different word embedding dimensions and semantic categories.
Experimental results
Research questions
- RQ1To what extent do word embeddings trained on web text reproduce known human semantic biases from psychological studies?
- RQ2Can standard NLP models like GloVe automatically learn and reflect societal biases present in language corpora?
- RQ3How do the associations between gender and professions in word embeddings compare to real-world demographic distributions?
- RQ4Can the WEAT and WEFAT frameworks reliably detect and quantify bias in word embeddings?
- RQ5Does mere exposure to natural language, without explicit instruction, lead to the internalization of social biases in machine learning models?
Key findings
- Word embeddings trained on web text replicate a wide spectrum of human-like biases, including those related to race, gender, and social roles, as measured by the WEAT.
- The study confirms that biases such as gender associations with careers (e.g., 'nurse' with 'female', 'engineer' with 'male') are accurately captured by GloVe models.
- The WEFAT test reveals that word embeddings reflect factual demographic distributions, such as gendered first names and occupational gender ratios, with high accuracy.
- The replication of biases is not due to model design but stems from the statistical regularities present in the training language data.
- The results show that even morally neutral associations—like those between 'flowers' and 'pleasant'—are encoded in embeddings, indicating that bias is a systemic feature of language-based AI.
- The findings suggest that language itself serves as a primary vector for embedding societal biases into machine learning systems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.