Skip to main content
QUICK REVIEW

[Paper Review] Dialect prejudice predicts AI decisions about people's character, employability, and criminality

Valentin Hofmann, Pratyusha Kalluri|arXiv (Cornell University)|Mar 1, 2024
Computational and Text Analysis Methods30 citations
TL;DR

The paper develops Matched Guise Probing to reveal covert dialect prejudice against African American English in multiple language models, and shows this bias influences employment and criminality judgments even when race is not overtly stated.

ABSTRACT

Hundreds of millions of people now interact with language models, with uses ranging from serving as a writing aid to informing hiring decisions. Yet these language models are known to perpetuate systematic racial prejudices, making their judgments biased in problematic ways about groups like African Americans. While prior research has focused on overt racism in language models, social scientists have argued that racism with a more subtle character has developed over time. It is unknown whether this covert racism manifests in language models. Here, we demonstrate that language models embody covert racism in the form of dialect prejudice: we extend research showing that Americans hold raciolinguistic stereotypes about speakers of African American English and find that language models have the same prejudice, exhibiting covert stereotypes that are more negative than any human stereotypes about African Americans ever experimentally recorded, although closest to the ones from before the civil rights movement. By contrast, the language models' overt stereotypes about African Americans are much more positive. We demonstrate that dialect prejudice has the potential for harmful consequences by asking language models to make hypothetical decisions about people, based only on how they speak. Language models are more likely to suggest that speakers of African American English be assigned less prestigious jobs, be convicted of crimes, and be sentenced to death. Finally, we show that existing methods for alleviating racial bias in language models such as human feedback training do not mitigate the dialect prejudice, but can exacerbate the discrepancy between covert and overt stereotypes, by teaching language models to superficially conceal the racism that they maintain on a deeper level. Our findings have far-reaching implications for the fair and safe employment of language technology.

Motivation & Objective

  • Investigate whether language models hold covert raciolinguistic stereotypes activated by dialect features rather than explicit race.
  • Develop and apply a probing method (Matched Guise Probing) to detect dialect prejudice across models and settings.
  • Assess how dialect prejudice affects AI decisions in employment and criminal justice contexts.
  • Evaluate whether common bias-mitigation strategies (scaling, human feedback) reduce covert dialect prejudice.

Proposed method

  • Introduce Matched Guise Probing to compare predictions for AAE vs. SAE texts without overt race mention.
  • Analyze multiple models (GPT2, RoBERTa, T5, GPT3.5, GPT4) across meaning-matched and non-meaning-matched prompts.
  • Measure covert stereotypes by ranking adjectives associated with AAE against human stereotypes from Princeton Trilogy studies.
  • Assess employability by matching occupations to speakers of AAE vs. SAE and examining prestige correlations.
  • Assess criminality by simulating trials and computing conviction and death-sentence rates for AAE vs. SAE utterances.
  • Examine scaling and human-feedback effects on overt vs. covert stereotypes.

Experimental results

Research questions

  • RQ1Do language models exhibit covert dialect prejudice triggered by AAE features, independent of explicit racial cues?
  • RQ2How do covert stereotypes compare to overt stereotypes in language models, and how do they align with historical human stereotypes?
  • RQ3Do dialect-based biases influence AI judgments in employment and criminal justice scenarios?
  • RQ4Can model scaling or human-feedback training mitigate covert dialect prejudice?

Key findings

  • Covert stereotypes about AAE in language models align with archaic human stereotypes from the 1930s, and are more negative than any experimentally recorded modern human stereotypes.
  • Overt stereotypes about African Americans in several models are positive, especially in models trained with human feedback, creating a mismatch between covert and overt biases.
  • In employment tasks, models associate AAE speech with lower prestige occupations and higher association with SAEs, predicting reduced occupational prestige for AAE speakers.
  • In criminality tasks, models exhibit higher conviction rates and death-penalty selections for AAE utterances compared with SAE utterances.
  • Model scaling increases covert dialect prejudice (despite improving understanding of AAE) and reduces overt prejudice; human-feedback training increases overt positivity but does not reduce covert prejudice.
  • Human feedback reduces overt stereotypes but leaves covert stereotypes largely unchanged, amplifying the covert-overta gap in some models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.