[Paper Review] Regional Negative Bias in Word Embeddings Predicts Racial Animus--but only via Name Frequency
This paper demonstrates that anti-Black bias estimates from the Word Embedding Association Test (WEAT) in U.S. metropolitan areas are not genuine linguistic biases but are instead artifacts driven by the relative frequency of Black names in social media data. When controlling for name frequency, all correlations between WEAT scores and measures of racial animus—such as implicit bias and segregation—disappear, revealing a spurious relationship rooted in term frequency rather than semantic bias.
The word embedding association test (WEAT) is an important method for measuring linguistic biases against social groups such as ethnic minorities in large text corpora. It does so by comparing the semantic relatedness of words prototypical of the groups (e.g., names unique to those groups) and attribute words (e.g., 'pleasant' and 'unpleasant' words). We show that anti-black WEAT estimates from geo-tagged social media data at the level of metropolitan statistical areas strongly correlate with several measures of racial animus--even when controlling for sociodemographic covariates. However, we also show that every one of these correlations is explained by a third variable: the frequency of Black names in the underlying corpora relative to White names. This occurs because word embeddings tend to group positive (negative) words and frequent (rare) words together in the estimated semantic space. As the frequency of Black names on social media is strongly correlated with Black Americans' prevalence in the population, this results in spurious anti-Black WEAT estimates wherever few Black Americans live. This suggests that research using the WEAT to measure bias should consider term frequency, and also demonstrates the potential consequences of using black-box models like word embeddings to study human cognition and behavior.
Motivation & Objective
- To investigate whether WEAT-based estimates of anti-Black linguistic bias in U.S. metropolitan areas reflect genuine racial animus or methodological artifacts.
- To examine whether the observed correlations between WEAT scores and regional racial animus measures are spurious due to unobserved variables.
- To test whether the relative frequency of Black names in social media corpora explains the apparent anti-Black bias in word embeddings.
- To evaluate the robustness of WEAT as a measure of linguistic bias when standard sociodemographic controls are applied.
- To highlight the risks of using black-box NLP models like word embeddings to study human cognition and social attitudes without accounting for confounding factors like term frequency.
Proposed method
- The study uses geo-tagged U.S. Twitter data to compute WEAT scores at the metropolitan statistical area (MSA) level, measuring anti-Black bias via cosine similarity between word embeddings of Black names, White names, and pleasant/unpleasant attribute words.
- It applies multiple regression models to test the relationship between WEAT scores and regional measures of racial animus, including implicit bias (IAT scores) and residential segregation.
- The key methodological innovation is the inclusion of relative Black name frequency—defined as the proportion of uniquely Black names in the corpus relative to White names—as a control variable.
- The authors compare model fit and coefficient significance across models with and without name frequency control to isolate its explanatory power.
- They use LOWESS smoothing to visualize the strong correlation between the proportion of Black residents and the proportion of Black names in the data.
- The analysis relies on Word2Vec embeddings trained on Twitter corpora, with WEAT scores calculated using the standard formula involving mean cosine similarities between attribute and category word pairs.
Experimental results
Research questions
- RQ1Does the WEAT detect genuine anti-Black linguistic bias in regional U.S. social media corpora, or is it confounded by demographic factors?
- RQ2To what extent does the relative frequency of Black names in a corpus explain the observed correlation between WEAT scores and measures of racial animus?
- RQ3Are standard sociodemographic controls sufficient to correct for the bias introduced by term frequency in word embeddings?
- RQ4Can the observed anti-Black WEAT scores be fully explained by the frequency of Black names, independent of actual racial attitudes?
- RQ5What are the implications of this confounding for the use of word embeddings in measuring cultural or social biases in computational social science?
Key findings
- The correlation between WEAT estimates and implicit bias (IAT scores) from Project Implicit is significant when unadjusted, but becomes non-significant after controlling for relative Black name frequency.
- The relationship between WEAT scores and residential segregation also vanishes when name frequency is included as a control variable.
- Controlling for name frequency reduces the magnitude of WEAT correlations with racial animus measures more than all standard sociodemographic controls combined.
- The proportion of Black names in the corpus explains over half of the variance in the observed anti-Black WEAT scores, indicating a strong confounding effect.
- The study finds that WEAT estimates are not measuring semantic bias but are instead a noisy proxy for the rarity of Black names in the data.
- The results suggest that word embeddings conflate term frequency with positivity, leading to spurious bias estimates when group representation varies across regions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.