Skip to main content
QUICK REVIEW

[Paper Review] Evaluation of African American Language Bias in Natural Language Generation

Nicholas Deas, Jessi Grieser|arXiv (Cornell University)|May 23, 2023
Text Readability and Simplification4 citations
TL;DR

This paper evaluates bias in large language models (LLMs) toward African American Language (AAL) using two tasks: counterpart generation and masked span prediction. Using a novel, multi-context AAL dataset, the authors find significant performance gaps, showing LLMs struggle more with AAL than White Mainstream English (WME), indicating dialectal bias in generation models.

ABSTRACT

We evaluate how well LLMs understand African American Language (AAL) in comparison to their performance on White Mainstream English (WME), the encouraged "standard" form of English taught in American classrooms. We measure LLM performance using automatic metrics and human judgments for two tasks: a counterpart generation task, where a model generates AAL (or WME) given WME (or AAL), and a masked span prediction (MSP) task, where models predict a phrase that was removed from their input. Our contributions include: (1) evaluation of six pre-trained, large language models on the two language generation tasks; (2) a novel dataset of AAL text from multiple contexts (social media, hip-hop lyrics, focus groups, and linguistic interviews) with human-annotated counterparts in WME; and (3) documentation of model performance gaps that suggest bias and identification of trends in lack of understanding of AAL features.

Motivation & Objective

  • To investigate whether large language models exhibit bias in understanding and generating African American Language (AAL) compared to White Mainstream English (WME).
  • To evaluate LLMs on two language generation tasks: counterpart generation and masked span prediction, using a novel, diverse AAL dataset.
  • To document performance gaps that suggest dialectal bias in LLMs, particularly in interpreting and producing AAL features such as dropped copulas and habitual be.
  • To raise awareness of ethical risks, including potential misuse in surveillance, while advocating for improved model inclusivity in high-impact applications like healthcare and mental health.

Proposed method

  • Constructed a novel dataset of AAL texts from social media, hip-hop lyrics, focus groups, and linguistic interviews, with human-annotated WME counterparts.
  • Evaluated six pre-trained LLMs on a counterpart generation task, where models translate between AAL and WME to assess dialectal understanding.
  • Applied a masked span prediction (MSP) task to test model ability to predict missing phrases in AAL and WME contexts.
  • Used automatic metrics and human judgments to assess model performance, focusing on accuracy and preservation of dialectal features.
  • Conducted error analysis to identify specific AAL morphosyntactic features—such as dropped copulas and aspect markers—where models consistently failed.
  • Employed a lexicon-based approach for toxicity evaluation instead of relying on potentially biased models to avoid reinforcing existing biases.

Experimental results

Research questions

  • RQ1Do large language models demonstrate performance gaps in generating and understanding African American Language (AAL) compared to White Mainstream English (WME)?
  • RQ2Which specific morphosyntactic features of AAL (e.g., dropped copulas, habitual be) are most frequently misinterpreted or misgenerated by LLMs?
  • RQ3How does model performance vary across different contexts of AAL use, such as social media, music, and focus groups?
  • RQ4To what extent do LLMs fail to preserve semantic equivalence when translating between AAL and WME?
  • RQ5What are the ethical implications of LLMs that better understand AAL, particularly in surveillance and high-stakes applications?

Key findings

  • LLMs exhibit significant performance gaps in understanding and generating AAL compared to WME, with consistent underperformance across all six models evaluated.
  • Models frequently fail to preserve key AAL features such as the dropped copula (e.g., 'she at work') and habitual be (e.g., 'he be running'), indicating poor dialectal comprehension.
  • Error analysis reveals that models are particularly prone to overcorrecting AAL features into WME forms, suggesting a bias toward standardization.
  • The masked span prediction task further confirms that models struggle to predict contextually appropriate AAL phrases, especially those involving non-standard grammar and lexical choices.
  • Performance gaps are more pronounced in complex, contextually nuanced AAL, such as in hip-hop lyrics and focus group discussions, where subtle linguistic cues carry semantic weight.
  • The study highlights that current evaluation metrics may themselves encode biases, as lexicon- and model-based measures can misclassify AAL as toxic or incorrect, reinforcing systemic inequities.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.