[Paper Review] ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over Centuries
ValNorm introduces a novel intrinsic evaluation task to quantify the valence (pleasantness/unpleasantness) dimension in word embeddings across seven languages and over 200 years of English text. It achieves high correlation (ρ = 0.88) with human-rated valence norms, revealing consistent, widely-shared affective associations in non-discriminatory words, while social group biases vary across languages and time.
Word embeddings learn implicit biases from linguistic regularities captured by word co-occurrence statistics. By extending methods that quantify human-like biases in word embeddings, we introduceValNorm, a novel intrinsic evaluation task and method to quantify the valence dimension of affect in human-rated word sets from social psychology. We apply ValNorm on static word embeddings from seven languages (Chinese, English, German, Polish, Portuguese, Spanish, and Turkish) and from historical English text spanning 200 years. ValNorm achieves consistently high accuracy in quantifying the valence of non-discriminatory, non-social group word sets. Specifically, ValNorm achieves a Pearson correlation of r=0.88 for human judgment scores of valence for 399 words collected to establish pleasantness norms in English. In contrast, we measure gender stereotypes using the same set of word embeddings and find that social biases vary across languages. Our results indicate that valence associations of non-discriminatory, non-social group words represent widely-shared associations, in seven languages and over 200 years.
Motivation & Objective
- To develop a method that quantifies the valence dimension of affect in word embeddings using human-rated psychological norms.
- To evaluate whether valence associations in word embeddings are consistent across languages and over time, independent of social group biases.
- To introduce ValNorm as a new intrinsic evaluation task that outperforms traditional word similarity metrics in capturing semantic quality.
- To distinguish between widely-shared, non-discriminatory valence norms and culture-specific or socially biased associations in word embeddings.
- To provide a transparent, statistically validated tool for assessing the semantic fidelity of word embeddings in capturing intrinsic affective meaning.
Proposed method
- ValNorm applies a permutation test to assess the statistical significance of valence quantification in word embeddings.
- It uses human-rated valence scores from Bellezza et al. (1986) as ground truth for 399 English words to compute correlation with embedding-based valence scores.
- The method computes a vector similarity score between target words and a set of positive/negative valence anchors (e.g., 'kindness' vs. 'torture') to estimate valence.
- ValNorm is applied to static word embeddings trained on diverse corpora and algorithms across seven languages and historical English text spanning 200 years.
- It compares performance against six traditional intrinsic evaluation tasks (e.g., word similarity) using Pearson correlation (ρ) as the metric.
- The approach isolates valence from syntactic and grammatical influences, such as gendered noun agreement, to ensure robustness to linguistic structure.
Experimental results
Research questions
- RQ1To what extent do word embeddings capture widely-shared, non-discriminatory valence norms across multiple languages?
- RQ2How stable are valence associations in word embeddings over a 200-year span of English historical text?
- RQ3How does ValNorm’s performance compare to traditional intrinsic evaluation tasks in measuring semantic quality?
- RQ4Are valence associations in word embeddings consistent across cultures and languages, or do they vary with linguistic structure?
- RQ5Can valence norms serve as a benchmark to detect shifted or harmful attitudes toward social groups in text corpora?
Key findings
- ValNorm achieves a Pearson correlation of ρ = 0.88 between embedding-based valence scores and human-rated valence norms for 399 English words, indicating high accuracy in capturing affective meaning.
- ValNorm maintains high performance across seven languages (Chinese, English, German, Polish, Portuguese, Spanish, Turkish), with ρ in the range [0.82, 0.88], demonstrating cross-linguistic consistency.
- Valence norms remain stable over 200 years of English text, with correlation coefficients ranging from ρ = 0.75 to 0.82 in historical embeddings, indicating long-term semantic stability.
- In contrast to valence norms, social group biases (e.g., gender, race) vary significantly across languages, suggesting cultural and linguistic dependence.
- ValNorm outperforms six traditional intrinsic evaluation tasks in measuring semantic quality, offering a more sensitive metric based on effect size rather than cosine similarity.
- The study confirms that non-discriminatory words like 'kindness' and 'vomit' maintain consistent pleasantness/unpleasantness associations across languages and time, reflecting widely-shared human affective responses.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.